Claude Opus 4.8 (xHigh) — benchmark results
Claude Opus 4.8 evaluated at the xhigh reasoning-effort setting. Provider: Anthropic. Released 2026-05-29. Access: API.
Unified ELO 1984 ± 26, rank #30 of 1776 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| WeirdML | 82.89 | Average Score | 96.4 |
| LiveBench | 77.55 | LiveBench average (self-reported) | 90.5 |
| LLM2014 Logic 2026-06 | 68.32 | Median Score | 90 |
| LLM2014 Logic 2026-07 | 66.7 | Median Score | 90 |
| LLM2014 Logic 2026-05 | 68.32 | Median Score | 89.5 |
| SWE-rebench | 56.55 | Resolved (%) | 85.6 |
| FrontierCode | 34.3 | Main Score (self-reported) | 83.3 |
| InferenceBench | 7.34 | Speedup Score | 81.8 |
| Creative Writing (Lechmazur) | 1.3 | Mean Score | 72.4 |
| Epoch AI - Cursorbench | 62.1 | Score | 70.4 |
| AutomationBench | 15.5 | Pass Rate (%) | 55.6 |
| Aikido CVE Rediscovery (pass@3) | 69.2 | Recall (%) | 45.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.