Claude Opus 4.8 (Max) — benchmark results
Claude Opus 4.8 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-05-29. Access: API.
Unified ELO 2002 ± 24, rank #23 of 1776 rated models, from 49 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MathArena - APEX 2025 | 81.25 | Accuracy (%) | 100 |
| CritPt | 20.9 | Accuracy (self-reported) | 97.9 |
| MathArena - Kangaroo 2025 Levels 11-12 | 100 | Accuracy (%) | 97.7 |
| MathArena - AIME 2026 | 100 | Accuracy (%) | 96.8 |
| OTIS Mock AIME 2024-25 | 98.33 | Accuracy (%) | 95.8 |
| Vals Index | 70.36 | Score (self-reported) | 95.8 |
| MathArena - Kangaroo 2025 Levels 7-8 | 96.67 | Accuracy (%) | 95.2 |
| FrontierMath - Tiers 1-3 | 47.24 | Accuracy (%, 290 problems) | 94.9 |
| Vals Multimodal Index | 70.89 | Score (self-reported) | 94.7 |
| Epoch AI - Apex Agents | 42.5 | Score | 93.8 |
| MathArena - ARXIV March | 75 | Accuracy (%) | 92.9 |
| MathArena - APEX Shortlist 2025 | 90.43 | Accuracy (%) | 91.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.