Claude Sonnet 5 (Max): benchmark results
Claude Sonnet 5 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-06-30. Access: API.
Unified ELO 1686 ± 1, rank #139 of 1761 rated models, from 43 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Vals AI MortgageTax | 70.03 | Accuracy (%) | 96.9 |
| Conceptual Reasoning Index - Argument Evaluation (LMCA) | 49.31 | Chance-Corrected Score (0-100) | 91.3 |
| Vals AI Vibe Code Bench | 81.33 | Accuracy (%) | 91.2 |
| Vals AI TaxEval v2 | 75.63 | Accuracy (%) | 90.6 |
| Vals AI CorpFin v2 | 66.98 | Accuracy (%) | 89.5 |
| Epoch AI - Mystery Game Puzzles | 35 | Score | 89.1 |
| Epoch AI - Dtbench | 92.48 | Score | 88.1 |
| ProofBench | 66 | Score (self-reported) | 87.7 |
| Vals AI ProgramBench | 72.07 | Raw Pass Rate (%) | 86.4 |
| Vals Index | 68.61 | Score (self-reported) | 84.6 |
| FrontierMath - Tiers 1-3 (v2) | 65.61 | Accuracy (%, 285 private v2 problems) | 83.5 |
| Vals AI (Vals Index) | 59.61 | Accuracy (%) | 83 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-max · How It Works · Data refreshed daily, snapshot 2026-09-05.