Claude Opus 4.6 (Thinking, High): benchmark results
Claude Opus 4.6 evaluated with thinking enabled at high reasoning effort. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1703 ± 1, rank #94 of 1761 rated models, from 31 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Generalization V2 (Lechmazur) | 80.6 | Inverse-Rank Score | 100 |
| Persuasion (Lechmazur) | 1.67 | Average Persuasion Strength | 92.9 |
| LiveBench Olympiad | 92.17 | Score | 91.5 |
| Buyout Game (Lechmazur) | 1759.7 | Bradley-Terry Rating | 88.6 |
| LiveBench | 76.79 | LiveBench average (self-reported) | 87.2 |
| Position Bias (Lechmazur) | 30.2 | Order Flip % (lower is better) | 85.7 |
| NYT Connections Extended | 88.1 | Score (%) | 82.7 |
| LiveBench Logic With Navigation | 80 | Score | 78.3 |
| LiveBench Code Generation | 80.28 | Score | 76.4 |
| LiveBench Theory of Mind | 82.69 | Score | 76.4 |
| LiveBench Plot Unscrambling | 66.48 | Score | 73.6 |
| LiveBench Typos | 84 | Score | 71.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-thinking-high · How It Works · Data refreshed daily, snapshot 2026-09-05.