Claude Opus 4.8 (Low): benchmark results
Provider: Anthropic. Released 2026-05-28. Access: API.
Unified ELO 1692 ± 1, rank #156 of 3081 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FinLifeBench - Financial State - Checkpoint State Accuracy | 80.1 | Cell-level state accuracy (%) | 100 |
| FinLifeBench - Financial State - Granular Change Accuracy | 47 | GCA@15 (%) | 100 |
| OTIS Mock AIME 2024-25 | 97.78 | Accuracy (%) | 93 |
| o11y-bench - Pass@3 | 87.3 | Tasks passed on at least one of three attempts, Pass@3 (%) | 85.3 |
| Epoch AI - GPQA Diamond | 88.38 | Accuracy (%) | 83.3 |
| ARC-AGI-2 | 62.22 | Accuracy (%) | 77.1 |
| Chess Puzzles (Epoch AI) | 29 | Accuracy (%) | 77.1 |
| ARC-AGI-1 | 88 | Accuracy (%) | 74.3 |
| o11y-bench - Pass^3 | 60.32 | Tasks passed on all three attempts, Pass^3 (%) | 71.6 |
| DataBench | 72 | Score (%) | 71 |
| DeepResearchBench | 49.3 | Average Score | 67.5 |
| ObviousBench | 95.14 | Answer pass³ (%) | 62.8 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-low · How It Works · Data refreshed daily, snapshot 2026-09-21.