Claude 3.5 Haiku: benchmark results

Anthropic's fast, low-cost tier of the Claude 3.5 generation, launched text-only with a 200K context (October 2024). Provider: Anthropic. Released 2024-11-04. Access: API.

Unified ELO 1530 ± 1, rank #543 of 1392 rated models, from 150 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LaborBench68.8F1 (self-reported)100
LLM Public Goods Game40.97Avg. Contribution (%)94.7
RAI-Bench - Refusal Rate (Career)100Rate (%)93.5
Arabic IFEval70.9Arabic Accuracy (%)92.3
Natural Language to Mongosh87.65NeXMaNeR (%)90.5
MERA - SimpleAr99.9EM (%)88.7
MERA - CheGeKa44.87F1 (%)87.6
RAI-Bench - Refusal Rate (General)82Rate (%)85.3
Bullshit Benchmark58.2BS Detection Rate (%)83
MERA - BPS99.1Accuracy (%)80.5
MERA - USE39.12Grade, normalized (%)80.4
AMA-Bench - Open-World QA61.38Average Score (%)79.2

Interactive version: theaggregate.ai/model?slug=claude-3-5-haiku · How It Works · Data refreshed daily, snapshot 2026-09-05.