Claude Opus 4.6 (Thinking, High) — benchmark results

Claude Opus 4.6 evaluated with thinking enabled at high reasoning effort. Provider: Anthropic. Released 2026-02-05. Access: API.

Unified ELO 1883 ± 39, rank #76 of 1776 rated models, from 8 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V2 (Lechmazur)80.6Inverse-Rank Score100
Persuasion (Lechmazur)1.67Average Persuasion Strength92.9
Buyout Game (Lechmazur)1759.7Bradley-Terry Rating88.6
NYT Connections Extended91.5Score (%)85.9
LiveBench76.79LiveBench average (self-reported)85.7
Position Bias (Lechmazur)30.2Order Flip % (lower is better)85.7
LLM Chess (Saplin)198.2ELO72.1
SkateBench64.36Success Rate (%)48.1

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-thinking-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.