Claude Opus 4.8: benchmark results

Anthropic Opus-tier Claude model for complex reasoning, coding, and agentic tasks. Provider: Anthropic. Released 2026-05-29. Access: API.

Unified ELO 1718 ± 1, rank #33 of 1392 rated models, from 654 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Agentic Skills Evaluation Framework92.7Overall Score w/ (self-reported)100
Benchmarks.bio - TxBench-PP59.33Pass Rate (%)100
BioMysteryBench Verified - Human Difficult (Anthropic)40Score (%)100
Bullshit Benchmark96.4BS Detection Rate (%)100
Business Utility Eval42Business Utility (%)100
CEO-Bench27776973Max final cash ($) (self-reported)100
ChartMuseum (Anthropic No Tools)75.8Accuracy (%)100
ChartMuseum (Anthropic Tools)89.7Accuracy (%)100
ChartQAPro (Anthropic No Tools)69.4Accuracy (%)100
ChartQAPro (Anthropic Tools)72.3Accuracy (%)100
Clerk LLM Leaderboard91.3Avg score (%)100
EdgeBench50.89Score @12h (134 tasks)100

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8 · How It Works · Data refreshed daily, snapshot 2026-09-05.