Claude Sonnet 4.5 (Thinking 16K): benchmark results
Claude Sonnet 4.5 evaluated with a 16K-token thinking budget. Provider: Anthropic. Released 2025-09-29. Access: API.
Unified ELO 1622 ± 1, rank #349 of 1761 rated models, from 8 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Elimination Game (Lechmazur) | 5.19 | TrueSkill μ | 83.1 |
| Step Game (Lechmazur) | 3.37 | TrueSkill μ | 81.1 |
| WeirdML | 47.71 | Average Score | 56.7 |
| OTIS Mock AIME 2024-25 | 71.11 | Accuracy (%) | 56.2 |
| ARC-AGI-2 | 6.94 | Accuracy (%) | 45.3 |
| ARC-AGI-1 | 48.33 | Accuracy (%) | 39.1 |
| NYT Connections Extended | 43 | Score (%) | 38.5 |
| IUMB | 8.3 | Score (%) | 3.7 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-5-thinking-16k · How It Works · Data refreshed daily, snapshot 2026-09-05.