Claude 3.7 Sonnet — benchmark results
Anthropic's first hybrid-reasoning Sonnet, switching between near-instant replies and extended thinking with an API thinking budget (February 2025). Provider: Anthropic. Released 2025-02-24. Access: API.
Unified ELO 1558 ± 7, rank #600 of 1776 rated models, from 718 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - riddlesense | 1.25 | Dataset z-score | 100 |
| DeepResearch Bench - Citation Accuracy | 87.32 | Score (%) | 100 |
| EmbodiedBench ALFRED | 67.7 | Avg Score (%) | 100 |
| OpenVLM MMT-Bench - DKR | 76.3 | Score (%) | 100 |
| OpenVLM MMT-Bench - Health and Medicine | 100 | Score (%) | 100 |
| PARROT - Oracle Dialect Compatibility | 58 | Accuracy (%) | 100 |
| WildVision-Bench | 60.6 | Overall Reward | 100 |
| OpenVLM MMT-Bench - 3D | 65 | Score (%) | 99.8 |
| OpenVLM MMT-Bench - Business | 91.7 | Score (%) | 99.8 |
| OpenVLM MMT-Bench - 3D Indoor Recognition | 60 | Score (%) | 99.5 |
| OpenVLM MMMU - Public Health | 96.7 | Accuracy (%) | 99.3 |
| OpenVLM MMT-Bench - Behavior Anomaly Detection | 80 | Score (%) | 99.3 |
Interactive version: theaggregate.ai/model?slug=claude-3-7-sonnet · How the rankings work · Data refreshed daily, snapshot 2026-07-22.