Claude Sonnet 4 (Thinking 16K) — benchmark results
Claude Sonnet 4 evaluated with a 16K-token thinking budget. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1748 ± 38, rank #202 of 1776 rated models, from 8 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Generalization V1 (Lechmazur) | 1.69 | Avg Rank (lower is better) | 98.8 |
| Confabulation Leaderboard (Lechmazur) | 2.48 | Confabulation rate % (lower is better) | 98.4 |
| NYT Connections Older Models | 40.3 | Score (%) | 69.4 |
| Step Game (Lechmazur) | 3.09 | TrueSkill μ | 68.2 |
| WeirdML | 46.11 | Average Score | 58.4 |
| Elimination Game (Lechmazur) | 4.32 | TrueSkill μ | 54.2 |
| ARC-AGI-2 | 5.93 | Accuracy (%) | 51.5 |
| ARC-AGI-1 | 40 | Accuracy (%) | 41 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-thinking-16k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.