Claude Sonnet 5.5 (Max): benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1954 ± 24, rank #18 of 2055 rated models, from 33 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Bug Hunt Bench | 57 | Planted Bugs Fixed (out of 105) | 100 |
| Bug Hunt Bench - LMS | 33 | Planted Bugs Fixed (out of 60) | 100 |
| Bug Hunt Bench - VS Code Extension | 24 | Planted Bugs Fixed (out of 45) | 100 |
| LiveBench Code Generation | 95.78 | Score | 100 |
| CursorBench 4.0 | 55.5 | Score (%) | 98.3 |
| Vals AI MedScribe | 91.1 | Accuracy (%) | 98.1 |
| LiveBench Code Completion | 86.96 | Score | 96.8 |
| LiveBench Theory of Mind | 86.54 | Score | 94.4 |
| LiveBench Olympiad | 92.47 | Score | 93.7 |
| LiveBench JavaScript | 77.27 | Score | 92.9 |
| LiveBench TypeScript | 56.67 | Score | 91.3 |
| Vals AI MedCode | 52.92 | Accuracy (%) | 90.1 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-5-max · How It Works · Data refreshed daily, snapshot 2026-09-29.