Claude Opus 4.8 (High) — benchmark results
Claude Opus 4.8 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2026-05-29. Access: API.
Unified ELO 1999 ± 38, rank #24 of 1776 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LisanBench | 0.66 | Mean Path Length / Current Maximum | 96.6 |
| ARC-AGI-2 | 72.08 | Accuracy (%) | 90.5 |
| GRIPS-hard | 76 | Accuracy (%) | 88.5 |
| Multi-turn Debate (Lechmazur) | 1668 | Bradley-Terry Rating | 87.5 |
| ARC-AGI-3 | 1.52 | Accuracy (%) | 85.7 |
| ARC-AGI-1 | 92 | Accuracy (%) | 85.4 |
| GRIPS | 93.7 | Accuracy (%) | 84.6 |
| PACT (Lechmazur) | 1562 | CMS Points | 84 |
| LiveBench | 76.18 | LiveBench average (self-reported) | 83.3 |
| Surface Evolver Bench | 87.5 | Mean Score (%) | 81.8 |
| Surface Evolver Bench Pass Rate | 68.75 | Pass Rate (%) | 79.5 |
| Vending-Bench 2 | 5787.43 | Money Balance ($) | 79.2 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.