Claude Opus 5 (Non-reasoning): benchmark results

Provider: Anthropic. Released 2026-07-24. Access: API.

Unified ELO 1694 ± 1, rank #160 of 3078 rated models, from 13 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
WeirdML86.34Average Score92.5
BBQ Ambiguous Accuracy (Sonnet 5 & Opus 5 System Cards)100Accuracy (%)90
PSF-Med - Frontier Hardest 5008.36Paraphrase Flip Rate (%)60
PSF-Med - Frontier Hardest 500 - VinDr-CXR10.6Paraphrase Flip Rate (%)60
Shade Coding Prompt Injection, Stronger Attacker - Attack Success Rate with Probes (Mythos 5.1 System Card)17Attack Success Rate (%)60
PSF-Med - Frontier Hardest 500 - PadChest5.79Paraphrase Flip Rate (%)40
Shade Coding Prompt Injection - Attack Success Rate (Mythos 5.1 System Card)3.61Attack Success Rate (%)40
Shade Coding Prompt Injection - Attack Success Rate with Probes (Mythos 5.1 System Card)0.42Attack Success Rate (%)40
Shade Coding Prompt Injection, Stronger Attacker - Attack Success Rate (Mythos 5.1 System Card)84.94Attack Success Rate (%)40
BBQ Disambiguated Accuracy (Sonnet 5 & Opus 5 System Cards)81.6Accuracy (%)25
PSF-Med - Frontier Hardest 500 - MIMIC-CXR9.3Paraphrase Flip Rate (%)20
Shade Computer Use Prompt Injection - Attack Success Rate (Mythos 5.1 System Card)3.57Attack Success Rate (%)20

Interactive version: theaggregate.ai/model?slug=claude-opus-5-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.