Claude Sonnet 4 (Thinking 16K) — benchmark results

Claude Sonnet 4 evaluated with a 16K-token thinking budget. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1748 ± 38, rank #202 of 1776 rated models, from 8 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V1 (Lechmazur)1.69Avg Rank (lower is better)98.8
Confabulation Leaderboard (Lechmazur)2.48Confabulation rate % (lower is better)98.4
NYT Connections Older Models40.3Score (%)69.4
Step Game (Lechmazur)3.09TrueSkill μ68.2
WeirdML46.11Average Score58.4
Elimination Game (Lechmazur)4.32TrueSkill μ54.2
ARC-AGI-25.93Accuracy (%)51.5
ARC-AGI-140Accuracy (%)41

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-thinking-16k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.