Claude Opus 4.8 (Thinking): benchmark results
Claude Opus 4.8 evaluated with thinking enabled. Provider: Anthropic. Released 2026-05-28. Access: API.
Unified ELO 1739 ± 1, rank #50 of 3078 rated models, from 38 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ComboShoppingBench - Response Quality | 96.6 | Pass rate (%; LLM-judged) | 100 |
| ReLE - Reasoning and Mathematics | 89.9 | Accuracy (%) | 99.4 |
| Kagi LLM Benchmark | 88.8 | Accuracy (%) | 98.6 |
| ComboShoppingBench - Coupon-ID Validity | 100 | Pass rate (%) | 97.6 |
| ReLE - Education - Primary School Subjects | 70.7 | Accuracy (%) | 97.2 |
| Wolfram LLM Benchmarking Project | 68.1 | Correct Functionality (%) | 96.3 |
| ReLE - Reasoning - BBH | 85.7 | Accuracy (%) | 94.4 |
| GeoBench Photos | 4101.62 | Average Score | 93.5 |
| ComboShoppingBench - Claim Faithfulness | 94.5 | Pass rate (%; LLM-judged) | 92.9 |
| ComboShoppingBench - Coupon Legality | 98.6 | Pass rate (%) | 92.9 |
| ReLE - Overall | 74.7 | Accuracy (%) | 91.9 |
| ProfBench | 56.8 | Overall Rubric Score (%) | 88.6 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.