Claude Sonnet 4 (Non-reasoning): benchmark results

Claude Sonnet 4 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1572 ± 1, rank #549 of 1761 rated models, from 32 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI - Writing63.68Writing Score94.7
UGI - Natural Intelligence51.13NatInt Score90.6
Confabulation Leaderboard (Lechmazur)5.45Confabulation rate % (lower is better)85.7
AA Omniscience-9.02Score74.1
AA Omniscience - Software Engineering (SWE)39.1Accuracy (%)70.8
Elimination Game (Lechmazur)4.64TrueSkill μ69.5
AA Terminal-Bench Hard27.27Accuracy (%)69.3
LLM Emergent Collusion32Collusion Rate (%)66.7
AA CritPt1.14Accuracy (%)65.5
Artificial Analysis Intelligence Index19.06Intelligence Index64.5
Aider polyglot coding leaderboard56.4Pass rate (%)60.6
Generalization V1 (Lechmazur)1.89Avg Rank (lower is better)60

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.