Claude Opus 4.6 (Max): benchmark results
Claude Opus 4.6 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1691 ± 1, rank #128 of 1761 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ATLAS | 70.55 | ATLAScore (self-reported) | 100 |
| Conceptual Reasoning Index - Argument Evaluation (LMCA) | 55.76 | Chance-Corrected Score (0-100) | 97.1 |
| VideoMME w sub. | 86.1 | Score (self-reported) | 92.7 |
| FrontierMath - Tiers 1-3 | 40.7 | Accuracy (%, 290 problems) | 92.4 |
| APEX-Agents | 48.4 | Mean Score (ReAct) (self-reported) | 92 |
| HMMT February 2026 | 96.2 | Score (self-reported) | 87.5 |
| FrontierMath - Tier 4 | 22.9 | Accuracy (%, 48 problems) | 87.3 |
| Vals AI MedScribe | 86.74 | Accuracy (%) | 86.6 |
| Epoch AI - Dtbench | 91.2 | Score | 85.9 |
| FrontierMath - Tiers 1-3 (v2) | 65.96 | Accuracy (%, 285 private v2 problems) | 84.6 |
| MathArena - APEX 2025 | 34.5 | Accuracy (%) | 80 |
| MedXpertQA | 64.4 | Score (self-reported) | 80 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-max · How It Works · Data refreshed daily, snapshot 2026-09-05.