Claude Opus 4.8 (Max) — benchmark results

Claude Opus 4.8 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-05-29. Access: API.

Unified ELO 2002 ± 24, rank #23 of 1776 rated models, from 49 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MathArena - APEX 202581.25Accuracy (%)100
CritPt20.9Accuracy (self-reported)97.9
MathArena - Kangaroo 2025 Levels 11-12100Accuracy (%)97.7
MathArena - AIME 2026100Accuracy (%)96.8
OTIS Mock AIME 2024-2598.33Accuracy (%)95.8
Vals Index70.36Score (self-reported)95.8
MathArena - Kangaroo 2025 Levels 7-896.67Accuracy (%)95.2
FrontierMath - Tiers 1-347.24Accuracy (%, 290 problems)94.9
Vals Multimodal Index70.89Score (self-reported)94.7
Epoch AI - Apex Agents42.5Score93.8
MathArena - ARXIV March75Accuracy (%)92.9
MathArena - APEX Shortlist 202590.43Accuracy (%)91.7

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.