Claude Opus 4.6 (Thinking 32K) — benchmark results

Claude Opus 4.6 evaluated with a 32K-token thinking budget. Provider: Anthropic. Released 2026-02-05. Access: API.

Unified ELO 1862 ± 74, rank #95 of 1776 rated models, from 7 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FrontierMath - Tiers 1-340Accuracy (%, 290 problems)89.9
Epoch AI - ECI155.38ECI Score86.7
FrontierMath - Tier 420.83Accuracy (%, 48 problems)85.2
OTIS Mock AIME 2024-2593.06Accuracy (%)84.6
SimpleQA Verified46.49Accuracy (%)64.1
NYT Connections Older Models28.1Score (%)55.6
Chess Puzzles (Epoch AI)17Accuracy (%)24.1

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-thinking-32k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.