Claude Sonnet 4.6 (Max) — benchmark results
Claude Sonnet 4.6 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-02-17. Access: API.
Unified ELO 1944 ± 40, rank #43 of 1776 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Finance Agent v1.1 | 63.33 | Score (self-reported) | 96.4 |
| AA-LCR | 70.7 | Score (self-reported) | 96.1 |
| CritPt | 3.1 | Accuracy (self-reported) | 86.4 |
| ARC-AGI-2 | 58.33 | Accuracy (%) | 79.8 |
| Vals Index | 60.06 | Score (self-reported) | 79.2 |
| Epoch AI - ECI | 153.18 | ECI Score | 78.7 |
| ITBench-AA | 39.8 | Average Precision at Full Recall (self-reported) | 77.3 |
| ARC-AGI-1 | 86 | Accuracy (%) | 76.1 |
| GENSTRAT | 64 | Alpha (chips/game) (self-reported) | 75 |
| Vals Multimodal Index | 60.57 | Score (self-reported) | 73.7 |
| Vals AI ProofBench | 45 | Accuracy (%) | 73.6 |
| APEX-Agents-AA | 28 | Pass@1 (self-reported) | 69.6 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.