Claude Sonnet 4.6 (Thinking, High) — benchmark results

Claude Sonnet 4.6 evaluated with thinking enabled at high reasoning effort. Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1829 ± 28, rank #115 of 1776 rated models, from 7 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V2 (Lechmazur)76.3Inverse-Rank Score86.2
Buyout Game (Lechmazur)1656.2Bradley-Terry Rating80
Persuasion (Lechmazur)1.58Average Persuasion Strength78.6
LLM Chess (Saplin)189.7ELO71.4
LiveBench75.59LiveBench average (self-reported)71.4
NYT Connections Extended80Score (%)68.5
Position Bias (Lechmazur)37.4Order Flip % (lower is better)60

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-thinking-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.