Claude Haiku 4.5 (Claude Code): benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1629 ± 26, rank #372 of 1605 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AgentEvalBench (Agent-Twostage)32.5Eval@1 (%; share of 40 cases, 20 agents under generic and sp100
QEncodeBench - Latin Square81.4Semantic pass@1 (%; L3 gate, Latin square completion (row an75
QEncodeBench49Semantic pass@1 (%; L3 gate on the 480-instance core set of 66.7
QEncodeBench - String Matching68.6Semantic pass@1 (%; L3 gate, string matching (the pattern oc66.7
QEncodeBench - Vertex Cover38.6Semantic pass@1 (%; L3 gate, vertex cover (edges covered wit66.7
QEncodeBench - 3-Coloring17Semantic pass@1 (%; L3 gate, graph 3-coloring (all edges bic58.3
QEncodeBench - 3-SAT46Semantic pass@1 (%; L3 gate, 3-SAT (all clauses satisfied), 58.3
SLDBench0.22Mean Reward R2 (score)50
QEncodeBench - Subset Sum57.1Semantic pass@1 (%; L3 gate, subset sum (selected values sum41.7
GameDevBench18.6Pass@1 (%)6.2
AgentEvalBench (Agent-Onestage)17.5Eval@1 (%; share of 40 cases, 20 agents under generic and sp0
AgentEvalBench (Agent-Sourcecode)45Eval@1 (%; share of 40 cases, 20 agents under generic and sp0

Interactive version: theaggregate.ai/model?slug=claude-haiku-4-5-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.