Claude Sonnet 4 (20250514) (Thinking) — benchmark results
Claude Sonnet 4 (20250514) evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-14. Access: API.
Unified ELO 1662 ± 20, rank #341 of 1776 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Writing | 62.31 | Writing Score | 94.6 |
| UGI - Natural Intelligence | 47.43 | NatInt Score | 89.8 |
| Web-Bench | 24.3 | Pass@1 (%) | 89.4 |
| Vals AI MATH 500 | 93.8 | Accuracy (%) | 73.9 |
| Vals AI MedQA | 92.71 | Accuracy (%) | 71.3 |
| Vals AI CorpFin v2 | 61.23 | Accuracy (%) | 61.5 |
| Vals AI LegalBench | 82.06 | Accuracy (%) | 56 |
| Vals AI TaxEval v2 | 72 | Accuracy (%) | 55.5 |
| Vals AI MMLU-Pro | 83.86 | Accuracy (%) | 55.4 |
| Vals AI MGSM | 90.87 | Accuracy (%) | 51.3 |
| Vals AI AIME | 76.25 | Accuracy (%) | 47.4 |
| Vals AI GPQA | 75 | Accuracy (%) | 44.3 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-20250514-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.