Claude Opus 4.7 (Max): benchmark results
Claude Opus 4.7 evaluated at the max reasoning-effort setting. Provider: Anthropic. Released 2026-04-16. Access: API.
Unified ELO 1712 ± 1, rank #80 of 1761 rated models, from 45 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HWE-Bench (Claude Code) | 74.6 | Resolved (%) | 100 |
| SuperCLUE-Writing - Content Quality | 90.77 | Score | 100 |
| SuperCLUE-Writing - Overall | 87.97 | Score | 100 |
| Vals AI SAGE | 56.1 | Accuracy (%) | 100 |
| Vals AI MortgageTax | 70.27 | Accuracy (%) | 97.9 |
| APEX-Agents | 50.6 | Mean Score (ReAct) (self-reported) | 95.5 |
| Vals AI MMLU-Pro | 89.87 | Accuracy (%) | 94.9 |
| Vals AI MedCode | 54.86 | Accuracy (%) | 94.4 |
| Conceptual Reasoning Index - Argument Evaluation (LMCA) | 52.2 | Chance-Corrected Score (0-100) | 93.6 |
| MedCode | 54.86 | Score (self-reported) | 93.2 |
| SuperCLUE-Writing - Constrained Writing | 95.41 | Score | 92.3 |
| Epoch AI - Dtbench | 94.65 | Score | 91.2 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-7-max · How It Works · Data refreshed daily, snapshot 2026-09-05.