Claude Sonnet 4.5 (Thinking 16K) — benchmark results
Claude Sonnet 4.5 evaluated with a 16K-token thinking budget. Provider: Anthropic. Released 2025-09-29. Access: API.
Unified ELO 1694 ± 56, rank #283 of 1776 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Elimination Game (Lechmazur) | 5.19 | TrueSkill μ | 83.1 |
| Step Game (Lechmazur) | 3.37 | TrueSkill μ | 81.1 |
| Epoch AI - ECI | 146.76 | ECI Score | 63.7 |
| WeirdML | 47.71 | Average Score | 62 |
| OTIS Mock AIME 2024-25 | 71.11 | Accuracy (%) | 56.1 |
| ARC-AGI-2 | 6.94 | Accuracy (%) | 54.9 |
| ARC-AGI-1 | 48.33 | Accuracy (%) | 47.8 |
| NYT Connections Extended | 54 | Score (%) | 42.4 |
| IUMB | 8.3 | Score (%) | 3.6 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-5-thinking-16k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.