Claude Sonnet 4.5 (Thinking 16K) — benchmark results

Claude Sonnet 4.5 evaluated with a 16K-token thinking budget. Provider: Anthropic. Released 2025-09-29. Access: API.

Unified ELO 1694 ± 56, rank #283 of 1776 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Elimination Game (Lechmazur)5.19TrueSkill μ83.1
Step Game (Lechmazur)3.37TrueSkill μ81.1
Epoch AI - ECI146.76ECI Score63.7
WeirdML47.71Average Score62
OTIS Mock AIME 2024-2571.11Accuracy (%)56.1
ARC-AGI-26.94Accuracy (%)54.9
ARC-AGI-148.33Accuracy (%)47.8
NYT Connections Extended54Score (%)42.4
IUMB8.3Score (%)3.6

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-5-thinking-16k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.