Claude 3.7 Sonnet (20250219) (Thinking) — benchmark results

Claude 3.7 Sonnet (20250219) evaluated with thinking enabled. Provider: Anthropic. Released 2025-02-19. Access: API.

Unified ELO 1626 ± 13, rank #421 of 1776 rated models, from 122 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Danish NLU - Dansk72.62Named entity recognition Score (%)100
EuroEval Norwegian NLU - ScaLA NB81.62Linguistic acceptability Score (%)100
EuroEval Swedish NLU77.77NLU Average Score (%)100
EuroEval Swedish NLU - ScaLA SV80.86Linguistic acceptability Score (%)100
EuroEval Danish78.33Average Score (%)99.8
EuroEval Icelandic NLU - ScaLA IS62.13Linguistic acceptability Score (%)99.7
EuroEval Danish NLU - ScaLA DA76.38Linguistic acceptability Score (%)99.5
EuroEval Faroese NLU - ScaLA FO49.81Linguistic acceptability Score (%)99.3
EuroEval Dutch NLU - ScaLA NL74.66Linguistic acceptability Score (%)98.8
EuroEval Danish NLU69.11NLU Average Score (%)98.6
EuroEval Norwegian73.12Average Score (%)98.6
EuroEval Faroese67.5Average Score (%)98.5

Interactive version: theaggregate.ai/model?slug=claude-3-7-sonnet-20250219-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.