Claude 3.5 Haiku — benchmark results

Anthropic's fast, low-cost tier of the Claude 3.5 generation, launched text-only with a 200K context (October 2024). Provider: Anthropic. Released 2024-11-04. Access: API.

Unified ELO 1501 ± 14, rank #794 of 1776 rated models, from 138 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LaborBench68.8F1 (self-reported)100
LLM Public Goods Game40.97Avg. Contribution (%)94.7
Arabic IFEval70.9Arabic Accuracy (%)92.3
Bullshit Benchmark58.2BS Detection Rate (%)84.9
YapBench401.2YapIndex (lower is better)79.6
AMA-Bench - Open-World QA61.38Average Score (%)78.3
Judge Arena1282ELO Score74.1
LLM Stats (DROP)83.1Score (%)73.2
Fin-Bias97.1Average Herding Score (with rating) (self-reported)72.2
AI for Education Visual Maths - Statistics and Probability28.57Accuracy (%)69.2
AA Omniscience-23.18Score68.8
BALROG NetHack (LLM)1.2Progress (%)66.7

Interactive version: theaggregate.ai/model?slug=claude-3-5-haiku · How the rankings work · Data refreshed daily, snapshot 2026-07-22.