Claude Opus 4.8 (xHigh): benchmark results
Claude Opus 4.8 evaluated at the xhigh reasoning-effort setting. Provider: Anthropic. Released 2026-05-29. Access: API.
Unified ELO 1713 ± 1, rank #77 of 1761 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ProgramBench Almost | 16.5 | Almost (%) | 95 |
| DuelLab Overall | 66.6 | DuelLab Score | 93.7 |
| WeirdML | 82.89 | Average Score | 92.4 |
| LiveBench | 77.55 | LiveBench average (self-reported) | 91.5 |
| EnigmaEval | 23.51 | Score (self-reported) | 90 |
| LLM2014 Logic 2026-06 | 68.32 | Median Score | 90 |
| LLM2014 Logic 2026-05 | 68.32 | Median Score | 89.5 |
| InferenceBench | 7.34 | Speedup Score | 82.3 |
| Epoch AI - Mystery Game Puzzles | 31 | Score | 81.5 |
| SWE-rebench | 56.55 | Resolved (%) | 79.2 |
| Creative Writing (Lechmazur) | 1.1 | Mean Score | 71.3 |
| CursorBench 3.1 | 59.4 | Score (%) | 52.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.