Claude Opus 4.6 (Thinking) — benchmark results
Claude Opus 4.6 evaluated with thinking enabled for harder reasoning tasks. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1944 ± 19, rank #41 of 1776 rated models, from 145 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ARC-AGI-2 Verified (Opus 4.6 System Card) | 68.8 | Score (%) | 100 |
| BioMysteryBench (Opus 4.6 System Card) | 61.5 | Score (%) | 100 |
| BioPipelineBench (Opus 4.6 System Card) | 53.1 | Score (%) | 100 |
| CharXiv Reasoning No Tools (Opus 4.6 System Card) | 68.5 | Score (%) | 100 |
| CharXiv Reasoning With Cropping (Opus 4.6 System Card) | 77.4 | Score (%) | 100 |
| Finance Agent (Opus 4.6 System Card) | 60.7 | Accuracy (%) | 100 |
| LAB-Bench FigQA With Cropping (Opus 4.6 System Card) | 78.3 | Accuracy (%) | 100 |
| LLM2014 Logic 2026-02 | 78.02 | Median Score | 100 |
| LLMEval-Logic Hard Sub-Q | 76.6 | Accuracy (%) | 100 |
| OSWorld (Opus 4.6 System Card) | 72.7 | First-Attempt Success Rate (%) | 100 |
| OpenRCA Banking (Opus 4.6 System Card) | 37.3 | Root Cause Identification (%) | 100 |
| OpenRCA Market (Opus 4.6 System Card) | 33.6 | Root Cause Identification (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.