Claude Sonnet 4 (2025-05-14) (Thinking): benchmark results
Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1625 ± 1, rank #359 of 1919 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety - Safety Scenarios | 98.07 | Mean score | 95.9 |
| HELM Safety | 98.07 | Mean score (self-reported) | 95 |
| HELM Safety - Harm Bench | 98 | LM Evaluated Safety score | 94.7 |
| HELM Capabilities - GPQA | 70.63 | COT correct | 90 |
| HELM Capabilities - Omni-MATH | 60.2 | Acc | 90 |
| HELM Capabilities - MMLU-Pro | 84.3 | COT correct | 89 |
| HELM Capabilities - WildBench | 83.81 | WB Score | 86 |
| HELM Safety - Simple Safety Tests | 100 | LM Evaluated Safety score | 84.6 |
| HELM Safety - Bbq | 96.6 | BBQ accuracy | 80.8 |
| HELM Capabilities - IFEval | 84.01 | IFEval Strict Acc | 68 |
| HELM Safety - Xstest | 96.61 | LM Evaluated Safety score | 63.9 |
| HELM Safety - Anthropic Red Team | 99.15 | LM Evaluated Safety score | 61.5 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-2025-05-14-thinking · How It Works · Data refreshed daily, snapshot 2026-09-08.