Claude Opus 4.6 (Low): benchmark results

Provider: Anthropic. Released 2026-02-05. Access: API.

Unified ELO 1660 ± 27, rank #481 of 2055 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Context Arena72.43Average Score (%)85.9
DeepResearchBench51.4Average Score82.5
GIM0.29IRT ability (theta)55.6
ObviousBench90.28Answer pass³ (%)47.9
NSMQ Riddles (Real-Time Proxy)52.56Exact match (%; 156 riddles from 2019, clues revealed seven 45.5
o11y-bench - Pass^349.21Tasks passed on all three attempts, Pass^3 (%)45.1
NSMQ Riddles82.05Exact match (%; 156 riddles from the 2019 NSMQ riddles round40
o11y-bench - Pass@369.84Tasks passed on at least one of three attempts, Pass@3 (%)27.5
NSMQ Riddles - Chemistry63.64Exact match (%; 44 chemistry riddles from the 2019 NSMQ ridd18.2
Epoch AI - Mystery Game Puzzles7Score11.4
BaFCo - Coarse Layout Analysis (CoT)1.31Mean average precision at IoU 0.3 (0-100): class-wise averag0
BaFCo - Coarse Layout Analysis (Zero-shot)1.68Mean average precision at IoU 0.3 (0-100): class-wise averag0

Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-low · How It Works · Data refreshed daily, snapshot 2026-09-29.