Claude Opus 4.8 (Max): benchmark results
Claude Opus 4.8 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-05-29. Access: API.
Unified ELO 1722 ± 1, rank #61 of 1761 rated models, from 105 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MathArena - APEX 2025 | 81.25 | Accuracy (%) | 100 |
| LiveBench Math Comp | 98.04 | Score | 99.1 |
| Vals AI Terminal-Bench 2.0 | 70.04 | Accuracy (%) | 98.5 |
| MathArena - Kangaroo 2025 Levels 11-12 | 100 | Accuracy (%) | 97.8 |
| Toolathlon | 79.9 | Score (self-reported) | 97.8 |
| Conceptual Reasoning Index - Argument Evaluation (LMCA) | 57.51 | Chance-Corrected Score (0-100) | 97.7 |
| Vals AI SAGE | 54.79 | Accuracy (%) | 97.5 |
| MathArena - AIME 2026 | 100 | Accuracy (%) | 96.8 |
| Vals AI MortgageTax | 69.91 | Accuracy (%) | 95.9 |
| LiveBench Code Completion | 84.78 | Score | 95.3 |
| OTIS Mock AIME 2024-25 | 98.33 | Accuracy (%) | 95.3 |
| FrontierMath - Tiers 1-3 | 47.24 | Accuracy (%, 290 problems) | 94.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-max · How It Works · Data refreshed daily, snapshot 2026-09-05.