Claude Opus 4.6 (Thinking) — benchmark results

Claude Opus 4.6 evaluated with thinking enabled for harder reasoning tasks. Provider: Anthropic. Released 2026-02-05. Access: API.

Unified ELO 1944 ± 19, rank #41 of 1776 rated models, from 145 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ARC-AGI-2 Verified (Opus 4.6 System Card)68.8Score (%)100
BioMysteryBench (Opus 4.6 System Card)61.5Score (%)100
BioPipelineBench (Opus 4.6 System Card)53.1Score (%)100
CharXiv Reasoning No Tools (Opus 4.6 System Card)68.5Score (%)100
CharXiv Reasoning With Cropping (Opus 4.6 System Card)77.4Score (%)100
Finance Agent (Opus 4.6 System Card)60.7Accuracy (%)100
LAB-Bench FigQA With Cropping (Opus 4.6 System Card)78.3Accuracy (%)100
LLM2014 Logic 2026-0278.02Median Score100
LLMEval-Logic Hard Sub-Q76.6Accuracy (%)100
OSWorld (Opus 4.6 System Card)72.7First-Attempt Success Rate (%)100
OpenRCA Banking (Opus 4.6 System Card)37.3Root Cause Identification (%)100
OpenRCA Market (Opus 4.6 System Card)33.6Root Cause Identification (%)100

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.