Claude Sonnet 4 (20250514) (Thinking) — benchmark results

Claude Sonnet 4 (20250514) evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-14. Access: API.

Unified ELO 1662 ± 20, rank #341 of 1776 rated models, from 22 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI - Writing62.31Writing Score94.6
UGI - Natural Intelligence47.43NatInt Score89.8
Web-Bench24.3Pass@1 (%)89.4
Vals AI MATH 50093.8Accuracy (%)73.9
Vals AI MedQA92.71Accuracy (%)71.3
Vals AI CorpFin v261.23Accuracy (%)61.5
Vals AI LegalBench82.06Accuracy (%)56
Vals AI TaxEval v272Accuracy (%)55.5
Vals AI MMLU-Pro83.86Accuracy (%)55.4
Vals AI MGSM90.87Accuracy (%)51.3
Vals AI AIME76.25Accuracy (%)47.4
Vals AI GPQA75Accuracy (%)44.3

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-20250514-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.