Claude Opus 5 (Adaptive Reasoning, Xhigh Effort): benchmark results
Provider: Anthropic. Released 2026-07-24. Access: API.
Unified ELO 1736 ± 1, rank #46 of 1761 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 53.37 | Intelligence Index | 99.2 |
| AA Humanity's Last Exam | 54.4 | Accuracy (%) | 98.8 |
| AA Omniscience - Science, Engineering & Mathematics | 56.51 | Accuracy (%) | 98.8 |
| AA Omniscience - Business | 51.5 | Accuracy (%) | 98.4 |
| AA GDPval | 1713.13 | ELO | 98.3 |
| AA Omniscience - Health | 52.42 | Accuracy (%) | 98.2 |
| AA GPQA Diamond | 93.74 | Accuracy (%) | 98 |
| AA Omniscience | 35.38 | Score | 97.9 |
| AA-Omniscience Accuracy | 59.5 | Accuracy (%) | 97.7 |
| AA Omniscience - Humanities & Social Sciences | 57.1 | Accuracy (%) | 97.3 |
| AA CritPt | 27.71 | Accuracy (%) | 96.6 |
| AA Omniscience - Software Engineering (SWE) | 83.7 | Accuracy (%) | 96 |
Interactive version: theaggregate.ai/model?slug=claude-opus-5-adaptive-reasoning-xhigh-effort · How It Works · Data refreshed daily, snapshot 2026-09-05.