Claude Haiku 4.5 (Claude Code): benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1629 ± 26, rank #372 of 1605 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AgentEvalBench (Agent-Twostage) | 32.5 | Eval@1 (%; share of 40 cases, 20 agents under generic and sp | 100 |
| QEncodeBench - Latin Square | 81.4 | Semantic pass@1 (%; L3 gate, Latin square completion (row an | 75 |
| QEncodeBench | 49 | Semantic pass@1 (%; L3 gate on the 480-instance core set of | 66.7 |
| QEncodeBench - String Matching | 68.6 | Semantic pass@1 (%; L3 gate, string matching (the pattern oc | 66.7 |
| QEncodeBench - Vertex Cover | 38.6 | Semantic pass@1 (%; L3 gate, vertex cover (edges covered wit | 66.7 |
| QEncodeBench - 3-Coloring | 17 | Semantic pass@1 (%; L3 gate, graph 3-coloring (all edges bic | 58.3 |
| QEncodeBench - 3-SAT | 46 | Semantic pass@1 (%; L3 gate, 3-SAT (all clauses satisfied), | 58.3 |
| SLDBench | 0.22 | Mean Reward R2 (score) | 50 |
| QEncodeBench - Subset Sum | 57.1 | Semantic pass@1 (%; L3 gate, subset sum (selected values sum | 41.7 |
| GameDevBench | 18.6 | Pass@1 (%) | 6.2 |
| AgentEvalBench (Agent-Onestage) | 17.5 | Eval@1 (%; share of 40 cases, 20 agents under generic and sp | 0 |
| AgentEvalBench (Agent-Sourcecode) | 45 | Eval@1 (%; share of 40 cases, 20 agents under generic and sp | 0 |
Interactive version: theaggregate.ai/model?slug=claude-haiku-4-5-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.