Claude Opus 4.8 (Max): benchmark results

Claude Opus 4.8 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-05-29. Access: API.

Unified ELO 1722 ± 1, rank #61 of 1761 rated models, from 105 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MathArena - APEX 202581.25Accuracy (%)100
LiveBench Math Comp98.04Score99.1
Vals AI Terminal-Bench 2.070.04Accuracy (%)98.5
MathArena - Kangaroo 2025 Levels 11-12100Accuracy (%)97.8
Toolathlon79.9Score (self-reported)97.8
Conceptual Reasoning Index - Argument Evaluation (LMCA)57.51Chance-Corrected Score (0-100)97.7
Vals AI SAGE54.79Accuracy (%)97.5
MathArena - AIME 2026100Accuracy (%)96.8
Vals AI MortgageTax69.91Accuracy (%)95.9
LiveBench Code Completion84.78Score95.3
OTIS Mock AIME 2024-2598.33Accuracy (%)95.3
FrontierMath - Tiers 1-347.24Accuracy (%, 290 problems)94.9

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-max · How It Works · Data refreshed daily, snapshot 2026-09-05.