Claude Opus 4.6 (Max) — benchmark results
Claude Opus 4.6 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1873 ± 31, rank #85 of 1776 rated models, from 44 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| APEX v1 Medicine (MD) | 71.6 | Score (%) | 100 |
| ATLAS | 70.55 | ATLAScore (self-reported) | 100 |
| AA-LCR | 70.7 | Score (self-reported) | 96.1 |
| CritPt | 12.6 | Accuracy (self-reported) | 95.7 |
| VideoMME w sub. | 86.1 | Score (self-reported) | 92.7 |
| FrontierMath - Tiers 1-3 | 40.7 | Accuracy (%, 290 problems) | 92.4 |
| MedScribe | 86.74 | Score (self-reported) | 92.3 |
| APEX v1 Big Law | 76.2 | Score (%) | 88.9 |
| Finance Agent v1.1 | 60.05 | Score (self-reported) | 87.3 |
| FrontierMath - Tier 4 | 22.9 | Accuracy (%, 48 problems) | 87.3 |
| APEX-Agents-AA | 33 | Pass@1 (self-reported) | 87 |
| HMMT February 2026 | 96.2 | Score (self-reported) | 87 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.