Claude Sonnet 4.6 (Thinking, High): benchmark results

Claude Sonnet 4.6 evaluated with thinking enabled at high reasoning effort. Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1670 ± 1, rank #180 of 1761 rated models, from 7 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V2 (Lechmazur)76.3Inverse-Rank Score86.2
Buyout Game (Lechmazur)1656.2Bradley-Terry Rating80
Persuasion (Lechmazur)1.58Average Persuasion Strength78.6
LiveBench75.59LiveBench average (self-reported)72.3
NYT Connections Extended74.1Score (%)66.3
LLM Chess (Saplin)189.7ELO63.4
Position Bias (Lechmazur)37.4Order Flip % (lower is better)60

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-thinking-high · How It Works · Data refreshed daily, snapshot 2026-09-05.