Claude Opus 4.6 (Thinking 32K) — benchmark results
Claude Opus 4.6 evaluated with a 32K-token thinking budget. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1862 ± 74, rank #95 of 1776 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FrontierMath - Tiers 1-3 | 40 | Accuracy (%, 290 problems) | 89.9 |
| Epoch AI - ECI | 155.38 | ECI Score | 86.7 |
| FrontierMath - Tier 4 | 20.83 | Accuracy (%, 48 problems) | 85.2 |
| OTIS Mock AIME 2024-25 | 93.06 | Accuracy (%) | 84.6 |
| SimpleQA Verified | 46.49 | Accuracy (%) | 64.1 |
| NYT Connections Older Models | 28.1 | Score (%) | 55.6 |
| Chess Puzzles (Epoch AI) | 17 | Accuracy (%) | 24.1 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-thinking-32k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.