Claude Opus 4.8 (Thinking): benchmark results

Claude Opus 4.8 evaluated with thinking enabled. Provider: Anthropic. Released 2026-05-28. Access: API.

Unified ELO 1739 ± 1, rank #50 of 3078 rated models, from 38 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ComboShoppingBench - Response Quality96.6Pass rate (%; LLM-judged)100
ReLE - Reasoning and Mathematics89.9Accuracy (%)99.4
Kagi LLM Benchmark88.8Accuracy (%)98.6
ComboShoppingBench - Coupon-ID Validity100Pass rate (%)97.6
ReLE - Education - Primary School Subjects70.7Accuracy (%)97.2
Wolfram LLM Benchmarking Project68.1Correct Functionality (%)96.3
ReLE - Reasoning - BBH85.7Accuracy (%)94.4
GeoBench Photos4101.62Average Score93.5
ComboShoppingBench - Claim Faithfulness94.5Pass rate (%; LLM-judged)92.9
ComboShoppingBench - Coupon Legality98.6Pass rate (%)92.9
ReLE - Overall74.7Accuracy (%)91.9
ProfBench56.8Overall Rubric Score (%)88.6

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.