Claude Opus 4.8 — benchmark results

Anthropic Opus-tier Claude model for complex reasoning, coding, and agentic tasks. Provider: Anthropic. Released 2026-05-29. Access: API.

Unified ELO 1901 ± 9, rank #63 of 1776 rated models, from 526 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Agentic Skills Evaluation Framework92.7Overall Score w/ (self-reported)100
ArXivMath Mar-Apr 2026 (Anthropic)71.82Accuracy (%)100
Benchmarks.bio - TxBench-PP59.33Pass Rate (%)100
BioMysteryBench Verified - Human Difficult (Anthropic)40Score (%)100
Bullshit Benchmark96.4BS Detection Rate (%)100
Business Utility Eval42Business Utility (%)100
CHI-Bench37.3Overall Pass@1 (%)100
ChartMuseum (Anthropic No Tools)75.8Accuracy (%)100
ChartMuseum (Anthropic Tools)89.7Accuracy (%)100
ChartQAPro (Anthropic No Tools)69.4Accuracy (%)100
ChartQAPro (Anthropic Tools)72.3Accuracy (%)100
Clerk LLM Leaderboard91.3Avg score (%)100

Interactive version: theaggregate.ai/model?slug=claude-opus-4-8 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.