Claude Opus 4.6 (Thinking, High): benchmark results

Claude Opus 4.6 evaluated with thinking enabled at high reasoning effort. Provider: Anthropic. Released 2026-02-05. Access: API.

Unified ELO 1703 ± 1, rank #94 of 1761 rated models, from 31 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V2 (Lechmazur)80.6Inverse-Rank Score100
Persuasion (Lechmazur)1.67Average Persuasion Strength92.9
LiveBench Olympiad92.17Score91.5
Buyout Game (Lechmazur)1759.7Bradley-Terry Rating88.6
LiveBench76.79LiveBench average (self-reported)87.2
Position Bias (Lechmazur)30.2Order Flip % (lower is better)85.7
NYT Connections Extended88.1Score (%)82.7
LiveBench Logic With Navigation80Score78.3
LiveBench Code Generation80.28Score76.4
LiveBench Theory of Mind82.69Score76.4
LiveBench Plot Unscrambling66.48Score73.6
LiveBench Typos84Score71.7

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-thinking-high · How It Works · Data refreshed daily, snapshot 2026-09-05.