Claude Sonnet 4 (Thinking): benchmark results
Claude Sonnet 4 evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1593 ± 1, rank #463 of 1761 rated models, from 101 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GAIA2 - Adaptability | 42.1 | Score (%) | 100 |
| GosuEvals | 74.3 | Score (%) | 100 |
| LLM2014 Code 2025-07 | 66.98 | Median Score | 100 |
| WildAgtEval | 67.5 | Complexity-Injected API Call Accuracy (self-reported) | 100 |
| SEAL - MASK | 95.33 | Score | 97 |
| LLM2014 Code 2025-09 | 61.01 | Median Score | 94.4 |
| GAIA2 | 37.8 | Pass@1 (%) | 93.3 |
| GAIA2 - Execution | 62.1 | Score (%) | 93.3 |
| GAIA2 - Noise | 31.2 | Score (%) | 93.3 |
| BenchTable | 74.7 | Total Score (%) | 93 |
| MCP-Universe (LLM w/ ReAct) | 51.74 | Avg Evaluator Score | 92.6 |
| LLM2014 Code 2025-09 - TypeScript | 8.89 | Score | 92.1 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.