Claude Opus 4.7 (xHigh): benchmark results

Claude Opus 4.7 evaluated at the xhigh reasoning-effort setting. Provider: Anthropic. Released 2026-04-16. Access: API.

Unified ELO 1712 ± 1, rank #79 of 1761 rated models, from 50 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LisanBench0.93Mean Path Length / Current Maximum99.3
LiveBench Math Comp98.04Score99.1
LiveBench Code Generation85.92Score96.2
OTIS Mock AIME 2024-2597.8Accuracy (%)94.8
FrontierMath - Tiers 1-343.8Accuracy (%, 290 problems)93.9
LiveBench77.1LiveBench average (self-reported)89.4
FrontierMath - Tier 422.92Accuracy (%, 48 problems)88.7
MathArena - APEX 202540.62Accuracy (%)85.5
LiveBench Spatial100Score84
LiveBench Consecutive Events89.52Score81.1
Chess Puzzles (Epoch AI)30Accuracy (%)79.9
ZeroBench15Score (%)78.3

Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.