Claude Opus 4.6 (High) — benchmark results
Claude Opus 4.6 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1860 ± 27, rank #96 of 1776 rated models, from 36 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CTI-REALM | 0.64 | Normalized Reward | 100 |
| DeepResearchBench | 55.31 | Average Score | 100 |
| WeirdML | 77.95 | Average Score | 94.9 |
| LLM2014 Logic 2026-05 | 76.48 | Median Score | 94.7 |
| APEX v1 Big Law | 76.4 | Score (%) | 94.4 |
| APEX v1 Medicine (MD) | 70.6 | Score (%) | 94.4 |
| MathArena - IMProofBench Final Answers | 80.74 | Accuracy (%) | 93.8 |
| MathArena - Project Euler 971-984 | 92.86 | Accuracy (%, direct Project Euler problems 971-984) | 90.9 |
| MathArena - ArXiv Math Dec 2025 | 57.35 | Accuracy (%) | 89.5 |
| MathArena - ArXiv Math Jan 2026 | 72.83 | Accuracy (%) | 88.5 |
| C4 Benchmark | 20.1 | Overall Score (%) | 87.5 |
| Epoch AI - ECI | 155.38 | ECI Score | 86.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.