Claude Sonnet 4 (20250514): benchmark results

May 14, 2025 Claude Sonnet 4 snapshot, tracked when sources report the dated API model. Provider: Anthropic. Released 2025-05-14. Access: API.

Unified ELO 1619 ± 1, rank #210 of 1392 rated models, from 113 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Galileo Agent - Healthcare TSQ95Task Success Quality (%)100
Galileo Agent - Healthcare Accuracy62Accuracy (%)97.6
Enkrypt AI - Safety Risk14.32Risk Score96.9
HELM AIR-Bench88.3Refusal Rate (%)96.5
HELM Safety BBQ97.9BBQ accuracy (%)94.8
Enkrypt AI - Toxicity Risk0.23Risk Score92.9
Galileo Agent - Banking Accuracy58Accuracy (%)92.9
Web-Bench25.1Pass@1 (%)92.6
HAL Online Mind2Web40Accuracy (%)90.9
RewardBench Factuality76.12Accuracy (%)90.8
RewardBench Math70.49Score (%)90.6
Galileo Agent - Insurance TSQ93Task Success Quality (%)90.5

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-20250514 · How It Works · Data refreshed daily, snapshot 2026-09-05.