Claude Sonnet 5 (Thinking): benchmark results

Provider: Anthropic. Released 2026-06-30. Access: API.

Unified ELO 1713 ± 1, rank #104 of 3078 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Anthropic Browser Use Prompt Injection - Attack Success Rate (Sonnet 5 & Opus 5 System Cards)0.93Attack Success Rate (%)100
Shade Coding Prompt Injection - Attack Success Rate (Sonnet 5 & Opus 5 System Cards)0.31Attack Success Rate (%)100
Shade Coding Prompt Injection - Attack Success Rate with Probes (Mythos 5.1 System Card)0.09Attack Success Rate (%)100
Shade Coding Prompt Injection - Attack Success Rate with Probes (Sonnet 5 & Opus 5 System Cards)0.09Attack Success Rate (%)100
Shade Coding Prompt Injection, Stronger Attacker - Attack Success Rate (Mythos 5.1 System Card)19.47Attack Success Rate (%)100
ProfBench60.4Overall Rubric Score (%)97.5
Sonar LLM Leaderboard - Java - Code Smell Density14.32Code smells per 1,000 lines of code (lower is better)97.2
ReLE - Reasoning - BBH85.7Accuracy (%)94.4
ReLE - Education - Primary School Subjects68.7Accuracy (%)89
Sonar LLM Leaderboard - Java - Pass Rate81.43Passing tests (%)86.8
ReLE - Reasoning and Mathematics82.4Accuracy (%)85.8
ReLE - Agents and Tool Use66.5Accuracy (%)84.8

Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.