Claude Opus 4.8 — benchmark results
Anthropic Opus-tier Claude model for complex reasoning, coding, and agentic tasks. Provider: Anthropic. Released 2026-05-29. Access: API.
Unified ELO 1901 ± 9, rank #63 of 1776 rated models, from 526 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Agentic Skills Evaluation Framework | 92.7 | Overall Score w/ (self-reported) | 100 |
| ArXivMath Mar-Apr 2026 (Anthropic) | 71.82 | Accuracy (%) | 100 |
| Benchmarks.bio - TxBench-PP | 59.33 | Pass Rate (%) | 100 |
| BioMysteryBench Verified - Human Difficult (Anthropic) | 40 | Score (%) | 100 |
| Bullshit Benchmark | 96.4 | BS Detection Rate (%) | 100 |
| Business Utility Eval | 42 | Business Utility (%) | 100 |
| CHI-Bench | 37.3 | Overall Pass@1 (%) | 100 |
| ChartMuseum (Anthropic No Tools) | 75.8 | Accuracy (%) | 100 |
| ChartMuseum (Anthropic Tools) | 89.7 | Accuracy (%) | 100 |
| ChartQAPro (Anthropic No Tools) | 69.4 | Accuracy (%) | 100 |
| ChartQAPro (Anthropic Tools) | 72.3 | Accuracy (%) | 100 |
| Clerk LLM Leaderboard | 91.3 | Avg score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-8 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.