Claude Sonnet 4 (Thinking) — benchmark results
Claude Sonnet 4 evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1704 ± 14, rank #264 of 1776 rated models, from 108 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GAIA2 - Adaptability | 42.1 | Score (%) | 100 |
| GosuEvals | 74.3 | Score (%) | 100 |
| LLM2014 Code 2025-07 | 66.98 | Median Score | 100 |
| AA MATH-500 | 99.07 | Accuracy (%) | 97.5 |
| SEAL - MASK | 95.33 | Score | 97 |
| LLM2014 Code 2025-09 | 61.01 | Median Score | 94.4 |
| GAIA2 | 37.8 | Pass@1 (%) | 93.3 |
| GAIA2 - Execution | 62.1 | Score (%) | 93.3 |
| GAIA2 - Noise | 31.2 | Score (%) | 93.3 |
| BenchTable | 74.7 | Total Score (%) | 93 |
| MCP-Universe (LLM w/ ReAct) | 51.74 | Avg Evaluator Score | 92.6 |
| MCP-Universe | 30.3 | Overall Success Rate (self-reported) | 92.3 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.