Claude Opus 5 (Low): benchmark results

Provider: Anthropic. Released 2026-07-24. Access: API.

Unified ELO 1733 ± 1, rank #62 of 3078 rated models, from 13 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Conceptual Reasoning Index - Consistency (ACCoRD)83.84Chance-Corrected Score (0-100)99.5
o11y-bench - Pass@392.06Tasks passed on at least one of three attempts, Pass@3 (%)98
Conceptual Reasoning Index - Argument Evaluation (LMCA)62.58Chance-Corrected Score (0-100)94.2
Conceptual Reasoning Index72.01Chance-Corrected Score (0-100)91.8
o11y-bench - Pass^368.25Tasks passed on all three attempts, Pass^3 (%)91.2
ObviousBench98.61Answer pass³ (%)86.9
Epoch AI - Critpt23.14Score84.6
OTIS Mock AIME 2024-2593.33Accuracy (%)83
Epoch AI - GPQA Diamond87.88Accuracy (%)81.4
Conceptual Reasoning Index - Decision Theory (DTBench)88.45Chance-Corrected Score (0-100)79.7
Epoch AI - Cursorbench62.8Score68.5
Chess Puzzles (Epoch AI)20Accuracy (%)60.6

Interactive version: theaggregate.ai/model?slug=claude-opus-5-low · How It Works · Data refreshed daily, snapshot 2026-09-19.