Claude 3.7 Sonnet — benchmark results

Anthropic's first hybrid-reasoning Sonnet, switching between near-instant replies and extended thinking with an API thinking budget (February 2025). Provider: Anthropic. Released 2025-02-24. Access: API.

Unified ELO 1558 ± 7, rank #600 of 1776 rated models, from 718 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AGC-Bench - riddlesense1.25Dataset z-score100
DeepResearch Bench - Citation Accuracy87.32Score (%)100
EmbodiedBench ALFRED67.7Avg Score (%)100
OpenVLM MMT-Bench - DKR76.3Score (%)100
OpenVLM MMT-Bench - Health and Medicine100Score (%)100
PARROT - Oracle Dialect Compatibility58Accuracy (%)100
WildVision-Bench60.6Overall Reward100
OpenVLM MMT-Bench - 3D65Score (%)99.8
OpenVLM MMT-Bench - Business91.7Score (%)99.8
OpenVLM MMT-Bench - 3D Indoor Recognition60Score (%)99.5
OpenVLM MMMU - Public Health96.7Accuracy (%)99.3
OpenVLM MMT-Bench - Behavior Anomaly Detection80Score (%)99.3

Interactive version: theaggregate.ai/model?slug=claude-3-7-sonnet · How the rankings work · Data refreshed daily, snapshot 2026-07-22.