Claude 3.7 Sonnet (20250219) (Thinking) — benchmark results
Claude 3.7 Sonnet (20250219) evaluated with thinking enabled. Provider: Anthropic. Released 2025-02-19. Access: API.
Unified ELO 1626 ± 13, rank #421 of 1776 rated models, from 122 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Danish NLU - Dansk | 72.62 | Named entity recognition Score (%) | 100 |
| EuroEval Norwegian NLU - ScaLA NB | 81.62 | Linguistic acceptability Score (%) | 100 |
| EuroEval Swedish NLU | 77.77 | NLU Average Score (%) | 100 |
| EuroEval Swedish NLU - ScaLA SV | 80.86 | Linguistic acceptability Score (%) | 100 |
| EuroEval Danish | 78.33 | Average Score (%) | 99.8 |
| EuroEval Icelandic NLU - ScaLA IS | 62.13 | Linguistic acceptability Score (%) | 99.7 |
| EuroEval Danish NLU - ScaLA DA | 76.38 | Linguistic acceptability Score (%) | 99.5 |
| EuroEval Faroese NLU - ScaLA FO | 49.81 | Linguistic acceptability Score (%) | 99.3 |
| EuroEval Dutch NLU - ScaLA NL | 74.66 | Linguistic acceptability Score (%) | 98.8 |
| EuroEval Danish NLU | 69.11 | NLU Average Score (%) | 98.6 |
| EuroEval Norwegian | 73.12 | Average Score (%) | 98.6 |
| EuroEval Faroese | 67.5 | Average Score (%) | 98.5 |
Interactive version: theaggregate.ai/model?slug=claude-3-7-sonnet-20250219-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.