Claude Sonnet 4 (20250514) — benchmark results

May 14, 2025 Claude Sonnet 4 snapshot, tracked when sources report the dated API model. Provider: Anthropic. Released 2025-05-14. Access: API.

Unified ELO 1673 ± 10, rank #314 of 1776 rated models, from 89 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Galileo Agent - Healthcare TSQ95Task Success Quality (%)100
Galileo Agent - Healthcare Accuracy62Accuracy (%)97.6
HELM AIR-Bench88.3Refusal Rate (%)96.5
HELM Safety BBQ97.9BBQ accuracy (%)94.8
Galileo Agent - Banking Accuracy58Accuracy (%)92.9
Web-Bench25.1Pass@1 (%)92.6
HAL Online Mind2Web40Accuracy (%)90.9
RewardBench Factuality76.12Accuracy (%)90.8
RewardBench Math70.49Score (%)90.6
Galileo Agent - Insurance TSQ93Task Success Quality (%)90.5
RewardBench Ties79.39Score (%)89.7
AraGen75.583C3H Score (%)89.6

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-20250514 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.