Claude Haiku 4.5 (20251001) — benchmark results

October 1, 2025 Claude Haiku 4.5 snapshot row. Provider: Anthropic. Released 2025-10-01. Access: API.

Unified ELO 1585 ± 7, rank #515 of 1776 rated models, from 303 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM AIR-Bench93.2Refusal Rate (%)100
EuroEval Spanish NLU - Sentiment Headlines ES52.68Sentiment classification Score (%)97.3
EuroEval Croatian NLU - MMS HR46.99Sentiment classification Score (%)97.2
EuroEval Greek Summarization - Greek Wikipedia32.52Score (%)94.5
EuroEval Spanish Common Sense Reasoning79.3Common Sense Reasoning Average Score (%)94.4
EuroEval Portuguese NLU - SST-2 PT85.2Sentiment classification Score (%)93.9
Gorilla API Bench (BFCL)68.7Overall Accuracy (%)93.9
HAL GAIA Level 361.54Accuracy (%)93.8
EuroEval French NLU - Allocine95.93Sentiment classification Score (%)92.3
EuroEval Faroese NLU - FoQA76.16Reading comprehension Score (%)92.1
EuroEval Dutch NLU - ScaLA NL63.67Linguistic acceptability Score (%)90.9
EuroEval English Common Sense Reasoning84.31Common Sense Reasoning Average Score (%)90.9

Interactive version: theaggregate.ai/model?slug=claude-haiku-4-5-20251001 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.