Claude Opus 4.7 (xHigh): benchmark results
Claude Opus 4.7 evaluated at the xhigh reasoning-effort setting. Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 1712 ± 1, rank #79 of 1761 rated models, from 50 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LisanBench | 0.93 | Mean Path Length / Current Maximum | 99.3 |
| LiveBench Math Comp | 98.04 | Score | 99.1 |
| LiveBench Code Generation | 85.92 | Score | 96.2 |
| OTIS Mock AIME 2024-25 | 97.8 | Accuracy (%) | 94.8 |
| FrontierMath - Tiers 1-3 | 43.8 | Accuracy (%, 290 problems) | 93.9 |
| LiveBench | 77.1 | LiveBench average (self-reported) | 89.4 |
| FrontierMath - Tier 4 | 22.92 | Accuracy (%, 48 problems) | 88.7 |
| MathArena - APEX 2025 | 40.62 | Accuracy (%) | 85.5 |
| LiveBench Spatial | 100 | Score | 84 |
| LiveBench Consecutive Events | 89.52 | Score | 81.1 |
| Chess Puzzles (Epoch AI) | 30 | Accuracy (%) | 79.9 |
| ZeroBench | 15 | Score (%) | 78.3 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.