Claude Sonnet 4 (Thinking 16K): benchmark results

Claude Sonnet 4 evaluated with a 16K-token thinking budget. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1628 ± 1, rank #329 of 1761 rated models, from 8 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V1 (Lechmazur)1.69Avg Rank (lower is better)98.8
Confabulation Leaderboard (Lechmazur)2.48Confabulation rate % (lower is better)98.4
NYT Connections Older Models26.4Score (%)68.2
Step Game (Lechmazur)3.09TrueSkill μ68.2
Elimination Game (Lechmazur)4.32TrueSkill μ54.2
WeirdML46.11Average Score52.9
ARC-AGI-25.93Accuracy (%)42.8
ARC-AGI-140Accuracy (%)34.1

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-thinking-16k · How It Works · Data refreshed daily, snapshot 2026-09-05.