Claude Sonnet 4 (Thinking) — benchmark results

Claude Sonnet 4 evaluated with thinking enabled. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1704 ± 14, rank #264 of 1776 rated models, from 108 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
GAIA2 - Adaptability42.1Score (%)100
GosuEvals74.3Score (%)100
LLM2014 Code 2025-0766.98Median Score100
AA MATH-50099.07Accuracy (%)97.5
SEAL - MASK95.33Score97
LLM2014 Code 2025-0961.01Median Score94.4
GAIA237.8Pass@1 (%)93.3
GAIA2 - Execution62.1Score (%)93.3
GAIA2 - Noise31.2Score (%)93.3
BenchTable74.7Total Score (%)93
MCP-Universe (LLM w/ ReAct)51.74Avg Evaluator Score92.6
MCP-Universe30.3Overall Success Rate (self-reported)92.3

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.