Claude 3.7 Sonnet (Thinking) — benchmark results
Claude 3.7 Sonnet evaluated with thinking enabled. Provider: Anthropic. Released 2025-02-24. Access: API.
Unified ELO 1652 ± 16, rank #362 of 1776 rated models, from 101 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| KORGym - Puzzle | 0.93 | Score | 94.4 |
| LLM2014 Logic 2025-03 | 73.76 | Median Score | 90 |
| BenchTable | 71 | Total Score (%) | 87.8 |
| AA Omniscience | 0.32 | Score | 86.8 |
| AA MMLU-Pro | 83.69 | Accuracy (%) | 86 |
| Gapminder AI Worldview | 87.9 | Correct Rate (%) | 85.3 |
| SEAL - Agentic Tool Use (Enterprise) | 65.27 | Score | 84.8 |
| AA Omniscience - Software Engineering (SWE) - Julia | 36 | Accuracy (%) | 84.6 |
| AidanBench | 2170 | Novel Answers | 83.6 |
| AA Omniscience - Software Engineering (SWE) - Dart | 38 | Accuracy (%) | 83.3 |
| AA Omniscience - Law | 26.1 | Accuracy (%) | 82.4 |
| LLM2014 Logic 2025-04 | 67.34 | Median Score | 82.1 |
Interactive version: theaggregate.ai/model?slug=claude-3-7-sonnet-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.