Claude Opus 4.8 (Adaptive Reasoning, Max Effort): benchmark results

Claude Opus 4.8 evaluated in adaptive-reasoning mode at max effort. Provider: Anthropic. Released 2026-05-29. Access: API.

Unified ELO 1675 ± 1, rank #164 of 1761 rated models, from 30 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Terminal-Bench Hard58.33Accuracy (%)97.9
AA Humanity's Last Exam48.66Accuracy (%)97.2
AA Omniscience28.75Score96.3
UGI - Writing65.88Writing Score96.2
UGI - Natural Intelligence65.39NatInt Score96.1
Artificial Analysis Intelligence Index46.44Intelligence Index94.9
AA Omniscience - Software Engineering (SWE)75.4Accuracy (%)94
UGI Leaderboard52.64UGI Score94
AA CritPt20.86Accuracy (%)93.3
AA GPQA Diamond92.02Accuracy (%)93.3
AA TAU-2 Bench94.44Accuracy (%)92.9
AA Omniscience - Health44.57Accuracy (%)92.4

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-adaptive-reasoning-max-effort · How It Works · Data refreshed daily, snapshot 2026-09-05.