Claude Opus 5.5 (xHigh): benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1785 ± 1, rank #10 of 3363 rated models, from 30 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| DataBench | 67 | Score (%) | 98.4 |
| LiveBench Code Generation | 91.55 | Score | 98.4 |
| ARC-AGI-2 | 92.5 | Accuracy (%) | 98.1 |
| LiveBench Code Completion | 86.96 | Score | 97.6 |
| ARC-AGI-1 | 97.5 | Accuracy (%) | 95.7 |
| LiveBench Integrals With Game | 99 | Score | 92.7 |
| Bug Hunt Bench | 36 | Planted Bugs Fixed (out of 105) | 92.4 |
| LiveBench AMPS Hard | 99 | Score | 91.1 |
| LiveBench Plot Unscrambling | 77.16 | Score | 90.3 |
| LiveBench Python | 70 | Score | 90.3 |
| Bug Hunt Bench - LMS | 18.5 | Planted Bugs Fixed (out of 60) | 89.9 |
| LiveBench Consecutive Events | 90.37 | Score | 88.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-5-5-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-23.