Claude Sonnet 5 (xHigh): benchmark results
Claude Sonnet 5 evaluated at the xhigh reasoning-effort setting. Provider: Anthropic. Released 2026-06-30. Access: API.
Unified ELO 1696 ± 1, rank #116 of 1761 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LiveBench Logic With Navigation | 86 | Score | 94.3 |
| LiveBench Code Generation | 83.1 | Score | 91.5 |
| OTIS Mock AIME 2024-25 | 94.72 | Accuracy (%) | 86.8 |
| Chess Puzzles (Epoch AI) | 35 | Accuracy (%) | 84.2 |
| LiveBench Spatial | 100 | Score | 84 |
| DuelLab Overall | 54.3 | DuelLab Score | 83.8 |
| LiveBench TypeScript | 50 | Score | 83 |
| LiveBench Python | 60 | Score | 77.4 |
| LiveBench Integrals With Game | 88 | Score | 72.6 |
| LiveBench JavaScript | 68.18 | Score | 70.8 |
| LLM2014 Logic 2026-07 | 43.14 | Median Score | 65.1 |
| LiveBench Theory of Mind | 80.77 | Score | 63.2 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.