Claude Opus 4.6 (High) — benchmark results

Claude Opus 4.6 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2026-02-05. Access: API.

Unified ELO 1860 ± 27, rank #96 of 1776 rated models, from 36 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CTI-REALM0.64Normalized Reward100
DeepResearchBench55.31Average Score100
WeirdML77.95Average Score94.9
LLM2014 Logic 2026-0576.48Median Score94.7
APEX v1 Big Law76.4Score (%)94.4
APEX v1 Medicine (MD)70.6Score (%)94.4
MathArena - IMProofBench Final Answers80.74Accuracy (%)93.8
MathArena - Project Euler 971-98492.86Accuracy (%, direct Project Euler problems 971-984)90.9
MathArena - ArXiv Math Dec 202557.35Accuracy (%)89.5
MathArena - ArXiv Math Jan 202672.83Accuracy (%)88.5
C4 Benchmark20.1Overall Score (%)87.5
Epoch AI - ECI155.38ECI Score86.7

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.