Claude Haiku 4.5 (Thinking) — benchmark results
Claude Haiku 4.5 evaluated with thinking enabled. Provider: Anthropic. Released 2025-10-01. Access: API.
Unified ELO 1665 ± 24, rank #332 of 1776 rated models, from 63 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Long Context Reasoning | 70.33 | Accuracy (%) | 93.7 |
| AA Omniscience | -4.22 | Score | 83.3 |
| AA SciCode | 43.29 | Accuracy (%) | 82.6 |
| AA AIME 2025 | 83.67 | Accuracy (%) | 80.4 |
| AA Omniscience - Software Engineering (SWE) - Julia | 28 | Accuracy (%) | 75.6 |
| Artificial Analysis Intelligence Index | 29.58 | Intelligence Index | 74.7 |
| BenchTable | 61.3 | Total Score (%) | 72.7 |
| AA Terminal-Bench Hard | 27.27 | Accuracy (%) | 69.3 |
| AA Omniscience - Software Engineering (SWE) - Rust | 58 | Accuracy (%) | 68.8 |
| AA Omniscience - Software Engineering (SWE) - HTML | 40 | Accuracy (%) | 67.5 |
| AA LiveCodeBench | 61.48 | Pass@1 (%) | 67 |
| YapBench | 335.2 | YapIndex (lower is better) | 66 |
Interactive version: theaggregate.ai/model?slug=claude-haiku-4-5-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.