Claude Haiku 4.5 — benchmark results
Anthropic Haiku-tier Claude model for fast, cost-efficient tasks. Provider: Anthropic. Released 2025-10-01. Access: API.
Unified ELO 1627 ± 9, rank #418 of 1776 rated models, from 529 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| DABstep | 89.95 | Hard Level Accuracy (%) | 100 |
| LiveSecBench | 91.43 | Overall Score (%) | 100 |
| MT-JailBench | 18.24 | CrescendoX ASR (self-reported) | 100 |
| RealityTest | 92.3 | Text disclosure probability (self-reported) | 100 |
| SWE-PRBench | 15.3 | Overall (sbar) (self-reported) | 100 |
| gwBenchmarks | 59.3 | Waveform (self-reported) | 100 |
| DystopiaBench | 23.8 | Dystopian Compliance Score (0-100) | 97.7 |
| Ko-AgentBench - L4 Parallel Tool Reasoning | 75 | Coverage (%) | 96.4 |
| Bullshit Benchmark | 87.3 | BS Detection Rate (%) | 95.8 |
| LLM-as-a-Reviewer | 99.7 | Low (NeurIPS 2022) (self-reported) | 95.5 |
| AA-LCR | 70.3 | Score (self-reported) | 95.3 |
| AGC-Bench - grapheval_ai_researcher | 1.06 | Dataset z-score | 93.8 |
Interactive version: theaggregate.ai/model?slug=claude-haiku-4-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.