Claude Opus 5 (Thinking): benchmark results

Provider: Anthropic. Released 2026-07-24. Access: API.

Unified ELO 1722 ± 1, rank #83 of 3078 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ProfBench66.2Overall Rubric Score (%)100
Shade Computer Use Prompt Injection - Attack Success Rate (Sonnet 5 & Opus 5 System Cards)0.54Attack Success Rate (%)100
Shade Computer Use Prompt Injection - Attack Success Rate with Probes (Sonnet 5 & Opus 5 System Cards)0.25Attack Success Rate (%)100
Sonar LLM Leaderboard - Java - Pass Rate88.6Passing tests (%)100
Sonar LLM Leaderboard - Java - Bug Density0.58Bugs per 1,000 lines of code (lower is better)93.8
Anthropic Browser Use Prompt Injection - Attack Success Rate (Sonnet 5 & Opus 5 System Cards)3.7Attack Success Rate (%)75
Shade Coding Prompt Injection - Attack Success Rate with Probes (Sonnet 5 & Opus 5 System Cards)0.18Attack Success Rate (%)75
PSF-Med - Frontier Hardest 500 - MIMIC-CXR7.88Paraphrase Flip Rate (%)60
Sonar LLM Leaderboard - Java - Vulnerability Density0.25Vulnerabilities per 1,000 lines of code (lower is better)57.6
Shade Coding Prompt Injection - Attack Success Rate (Sonnet 5 & Opus 5 System Cards)0.56Attack Success Rate (%)50
Sonar LLM Leaderboard - Java - Code Smell Density19.69Code smells per 1,000 lines of code (lower is better)47.2
PSF-Med - Frontier Hardest 5009.53Paraphrase Flip Rate (%)40

Interactive version: theaggregate.ai/model?slug=claude-opus-5-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.