Claude Opus 4.8 (Thinking) — benchmark results
Claude Opus 4.8 evaluated with thinking enabled. Provider: Anthropic. Released 2026-05-29. Access: API.
Unified ELO 1916 ± 26, rank #53 of 1776 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Kagi LLM Benchmark | 88.8 | Accuracy (%) | 98.6 |
| Wolfram LLM Benchmarking Project | 68.1 | Correct Functionality (%) | 97.2 |
| Chatbot Arena (Text) | 1484 | Elo | 96.8 |
| Chatbot Arena (Code) | 1565 | Elo | 96 |
| Agent Arena - Tool Hallucination | 0.22 | Tool Hallucination (%) | 94.6 |
| GeoBench Photos | 4101.62 | Average Score | 93.8 |
| Chatbot Arena (Vision) | 1286 | Arena Score | 93.3 |
| ProfBench | 56.8 | Overall Rubric Score (%) | 92 |
| Agent Arena - Praise vs Complaint | 19.42 | Praise vs Complaint (%) | 91.9 |
| GeoBench ACW | 4127.75 | Average Score | 87.2 |
| PM-LLM-Benchmark | 35.3 | Score | 86.2 |
| ProphetArena | 0.95 | 1 - Brier Score | 85.4 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.