Claude Opus 4.6 (Low): benchmark results
Provider: Anthropic. Released 2026-02-05. Access: API.
Unified ELO 1660 ± 27, rank #481 of 2055 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Context Arena | 72.43 | Average Score (%) | 85.9 |
| DeepResearchBench | 51.4 | Average Score | 82.5 |
| GIM | 0.29 | IRT ability (theta) | 55.6 |
| ObviousBench | 90.28 | Answer pass³ (%) | 47.9 |
| NSMQ Riddles (Real-Time Proxy) | 52.56 | Exact match (%; 156 riddles from 2019, clues revealed seven | 45.5 |
| o11y-bench - Pass^3 | 49.21 | Tasks passed on all three attempts, Pass^3 (%) | 45.1 |
| NSMQ Riddles | 82.05 | Exact match (%; 156 riddles from the 2019 NSMQ riddles round | 40 |
| o11y-bench - Pass@3 | 69.84 | Tasks passed on at least one of three attempts, Pass@3 (%) | 27.5 |
| NSMQ Riddles - Chemistry | 63.64 | Exact match (%; 44 chemistry riddles from the 2019 NSMQ ridd | 18.2 |
| Epoch AI - Mystery Game Puzzles | 7 | Score | 11.4 |
| BaFCo - Coarse Layout Analysis (CoT) | 1.31 | Mean average precision at IoU 0.3 (0-100): class-wise averag | 0 |
| BaFCo - Coarse Layout Analysis (Zero-shot) | 1.68 | Mean average precision at IoU 0.3 (0-100): class-wise averag | 0 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-6-low · How It Works · Data refreshed daily, snapshot 2026-09-29.