Claude Opus 4.8 (Low): benchmark results

Provider: Anthropic. Released 2026-05-28. Access: API.

Unified ELO 1692 ± 1, rank #156 of 3081 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FinLifeBench - Financial State - Checkpoint State Accuracy80.1Cell-level state accuracy (%)100
FinLifeBench - Financial State - Granular Change Accuracy47GCA@15 (%)100
OTIS Mock AIME 2024-2597.78Accuracy (%)93
o11y-bench - Pass@387.3Tasks passed on at least one of three attempts, Pass@3 (%)85.3
Epoch AI - GPQA Diamond88.38Accuracy (%)83.3
ARC-AGI-262.22Accuracy (%)77.1
Chess Puzzles (Epoch AI)29Accuracy (%)77.1
ARC-AGI-188Accuracy (%)74.3
o11y-bench - Pass^360.32Tasks passed on all three attempts, Pass^3 (%)71.6
DataBench72Score (%)71
DeepResearchBench49.3Average Score67.5
ObviousBench95.14Answer pass³ (%)62.8

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-low · How It Works · Data refreshed daily, snapshot 2026-09-21.