Claude Opus 4.7 (xHigh) — benchmark results

Claude Opus 4.7 evaluated at the xhigh reasoning-effort setting. Provider: Anthropic. Released 2026-04-16. Access: API.

Unified ELO 1963 ± 65, rank #38 of 1776 rated models, from 26 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LisanBench1Mean Path Length / Current Maximum100
OTIS Mock AIME 2024-2597.8Accuracy (%)94.9
FrontierMath - Tiers 1-343.79Accuracy (%, 290 problems)93.9
SWE-Milestone41.29Milestone Score (%)90.5
Epoch AI - ECI156.1ECI Score89.6
FrontierMath - Tier 422.92Accuracy (%, 48 problems)88.7
LiveBench77.1LiveBench average (self-reported)88.1
MathArena - APEX 202540.62Accuracy (%)86.4
ZeroBench14Score (%)86.4
ProgramBench Almost4.5Almost (%)83.3
SimpleQA Verified50.6Accuracy (%)74.2
MathArena - ARXIV April58.54Accuracy (%)71.4

Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.