Claude 3.7 Sonnet (Thinking): benchmark results
Claude 3.7 Sonnet evaluated with thinking enabled. Provider: Anthropic. Released 2025-02-24. Access: API.
Unified ELO 1585 ± 1, rank #497 of 1761 rated models, from 77 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CritPt | 90 | Accuracy (self-reported) | 97.7 |
| KORGym - Puzzle | 0.93 | Score | 94.4 |
| LLM2014 Logic 2025-03 | 73.76 | Median Score | 90 |
| BenchTable | 71 | Total Score (%) | 87.8 |
| Gapminder AI Worldview | 87.9 | Correct Rate (%) | 85.3 |
| SEAL - Agentic Tool Use (Enterprise) | 65.27 | Score | 84.8 |
| AidanBench | 2170 | Novel Answers | 83.6 |
| LLM2014 Logic 2025-04 | 67.34 | Median Score | 82.1 |
| GosuEvals | 71 | Score (%) | 81.5 |
| REAL Evals | 24 | Task Completion (%) | 80.4 |
| AA Omniscience | -0.73 | Score | 80.3 |
| Bullshit Benchmark | 52.7 | BS Detection Rate (%) | 79.4 |
Interactive version: theaggregate.ai/model?slug=claude-3-7-sonnet-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.