Claude Sonnet 4 (Non-reasoning): benchmark results
Claude Sonnet 4 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1572 ± 1, rank #549 of 1761 rated models, from 32 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Writing | 63.68 | Writing Score | 94.7 |
| UGI - Natural Intelligence | 51.13 | NatInt Score | 90.6 |
| Confabulation Leaderboard (Lechmazur) | 5.45 | Confabulation rate % (lower is better) | 85.7 |
| AA Omniscience | -9.02 | Score | 74.1 |
| AA Omniscience - Software Engineering (SWE) | 39.1 | Accuracy (%) | 70.8 |
| Elimination Game (Lechmazur) | 4.64 | TrueSkill μ | 69.5 |
| AA Terminal-Bench Hard | 27.27 | Accuracy (%) | 69.3 |
| LLM Emergent Collusion | 32 | Collusion Rate (%) | 66.7 |
| AA CritPt | 1.14 | Accuracy (%) | 65.5 |
| Artificial Analysis Intelligence Index | 19.06 | Intelligence Index | 64.5 |
| Aider polyglot coding leaderboard | 56.4 | Pass rate (%) | 60.6 |
| Generalization V1 (Lechmazur) | 1.89 | Avg Rank (lower is better) | 60 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.