Claude Opus 5 (Low): benchmark results
Provider: Anthropic. Released 2026-07-24. Access: API.
Unified ELO 1733 ± 1, rank #62 of 3078 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Conceptual Reasoning Index - Consistency (ACCoRD) | 83.84 | Chance-Corrected Score (0-100) | 99.5 |
| o11y-bench - Pass@3 | 92.06 | Tasks passed on at least one of three attempts, Pass@3 (%) | 98 |
| Conceptual Reasoning Index - Argument Evaluation (LMCA) | 62.58 | Chance-Corrected Score (0-100) | 94.2 |
| Conceptual Reasoning Index | 72.01 | Chance-Corrected Score (0-100) | 91.8 |
| o11y-bench - Pass^3 | 68.25 | Tasks passed on all three attempts, Pass^3 (%) | 91.2 |
| ObviousBench | 98.61 | Answer pass³ (%) | 86.9 |
| Epoch AI - Critpt | 23.14 | Score | 84.6 |
| OTIS Mock AIME 2024-25 | 93.33 | Accuracy (%) | 83 |
| Epoch AI - GPQA Diamond | 87.88 | Accuracy (%) | 81.4 |
| Conceptual Reasoning Index - Decision Theory (DTBench) | 88.45 | Chance-Corrected Score (0-100) | 79.7 |
| Epoch AI - Cursorbench | 62.8 | Score | 68.5 |
| Chess Puzzles (Epoch AI) | 20 | Accuracy (%) | 60.6 |
Interactive version: theaggregate.ai/model?slug=claude-opus-5-low · How It Works · Data refreshed daily, snapshot 2026-09-19.