Claude 3.5 Haiku — benchmark results
Anthropic's fast, low-cost tier of the Claude 3.5 generation, launched text-only with a 200K context (October 2024). Provider: Anthropic. Released 2024-11-04. Access: API.
Unified ELO 1501 ± 14, rank #794 of 1776 rated models, from 138 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LaborBench | 68.8 | F1 (self-reported) | 100 |
| LLM Public Goods Game | 40.97 | Avg. Contribution (%) | 94.7 |
| Arabic IFEval | 70.9 | Arabic Accuracy (%) | 92.3 |
| Bullshit Benchmark | 58.2 | BS Detection Rate (%) | 84.9 |
| YapBench | 401.2 | YapIndex (lower is better) | 79.6 |
| AMA-Bench - Open-World QA | 61.38 | Average Score (%) | 78.3 |
| Judge Arena | 1282 | ELO Score | 74.1 |
| LLM Stats (DROP) | 83.1 | Score (%) | 73.2 |
| Fin-Bias | 97.1 | Average Herding Score (with rating) (self-reported) | 72.2 |
| AI for Education Visual Maths - Statistics and Probability | 28.57 | Accuracy (%) | 69.2 |
| AA Omniscience | -23.18 | Score | 68.8 |
| BALROG NetHack (LLM) | 1.2 | Progress (%) | 66.7 |
Interactive version: theaggregate.ai/model?slug=claude-3-5-haiku · How the rankings work · Data refreshed daily, snapshot 2026-07-22.