Claude Opus 4.8 (Medium): benchmark results
Provider: Anthropic. Released 2026-05-28. Access: API.
Unified ELO 1681 ± 1, rank #201 of 3078 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| NOLLI - Average | 78.3 | Accuracy (%; 25-task macro average) | 92.9 |
| NOLLI - Cipher - English | 93 | Accuracy (%) | 92.9 |
| NOLLI - Cipher - Korean | 69.3 | Accuracy (%; jamo script) | 92.9 |
| NOLLI - Cryptarithmetic - English | 89.7 | Accuracy (%) | 92.9 |
| NOLLI - Cryptarithmetic - Korean | 95 | Accuracy (%; jamo script) | 92.9 |
| NOLLI - Direct Translations - English | 79.6 | Accuracy (%; eight translated puzzle types) | 92.9 |
| NOLLI - Direct Translations - Korean | 77.4 | Accuracy (%; eight translated puzzle types) | 92.9 |
| NOLLI - Korean Cultural Systems | 77.2 | Accuracy (%; kinship, saju, time and units) | 92.9 |
| NOLLI - Jamo Composition | 47 | Accuracy (%) | 85.7 |
| WeirdML | 76.04 | Average Score | 85.6 |
| ARC-AGI-2 | 71.67 | Accuracy (%) | 83.9 |
| ARC-AGI-1 | 91.5 | Accuracy (%) | 80.5 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-medium · How It Works · Data refreshed daily, snapshot 2026-09-19.