Claude Opus 4.8 (Medium): benchmark results

Provider: Anthropic. Released 2026-05-28. Access: API.

Unified ELO 1681 ± 1, rank #201 of 3078 rated models, from 21 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
NOLLI - Average78.3Accuracy (%; 25-task macro average)92.9
NOLLI - Cipher - English93Accuracy (%)92.9
NOLLI - Cipher - Korean69.3Accuracy (%; jamo script)92.9
NOLLI - Cryptarithmetic - English89.7Accuracy (%)92.9
NOLLI - Cryptarithmetic - Korean95Accuracy (%; jamo script)92.9
NOLLI - Direct Translations - English79.6Accuracy (%; eight translated puzzle types)92.9
NOLLI - Direct Translations - Korean77.4Accuracy (%; eight translated puzzle types)92.9
NOLLI - Korean Cultural Systems77.2Accuracy (%; kinship, saju, time and units)92.9
NOLLI - Jamo Composition47Accuracy (%)85.7
WeirdML76.04Average Score85.6
ARC-AGI-271.67Accuracy (%)83.9
ARC-AGI-191.5Accuracy (%)80.5

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-medium · How It Works · Data refreshed daily, snapshot 2026-09-19.