Claude Opus 4.7 (Max) — benchmark results
Claude Opus 4.7 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 2007 ± 42, rank #21 of 1776 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ITBench-AA | 46.7 | Average Precision at Full Recall (self-reported) | 100 |
| CritPt | 12 | Accuracy (self-reported) | 95.4 |
| AA-LCR | 70.3 | Score (self-reported) | 95.3 |
| MedCode | 54.86 | Score (self-reported) | 92.5 |
| WeirdML | 75.45 | Average Score | 90.5 |
| Epoch AI - ECI | 156.1 | ECI Score | 89.6 |
| Epoch AI - Scicode | 54.51 | Score | 85.3 |
| Epoch AI - Cursorbench | 64.8 | Score | 85.2 |
| IOI | 47.08 | Score (self-reported) | 84.9 |
| Epoch AI - Apex Agents | 33.9 | Score | 83.3 |
| Vals AI ProofBench | 54 | Accuracy (%) | 81.5 |
| MedScribe | 82.95 | Score (self-reported) | 75 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.