Claude Opus 4.6 (Thinking, High) — benchmark results
Claude Opus 4.6 evaluated with thinking enabled at high reasoning effort. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1883 ± 39, rank #76 of 1776 rated models, from 8 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Generalization V2 (Lechmazur) | 80.6 | Inverse-Rank Score | 100 |
| Persuasion (Lechmazur) | 1.67 | Average Persuasion Strength | 92.9 |
| Buyout Game (Lechmazur) | 1759.7 | Bradley-Terry Rating | 88.6 |
| NYT Connections Extended | 91.5 | Score (%) | 85.9 |
| LiveBench | 76.79 | LiveBench average (self-reported) | 85.7 |
| Position Bias (Lechmazur) | 30.2 | Order Flip % (lower is better) | 85.7 |
| LLM Chess (Saplin) | 198.2 | ELO | 72.1 |
| SkateBench | 64.36 | Success Rate (%) | 48.1 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-thinking-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.