Claude Sonnet 4.5 (Thinking) — benchmark results
Claude Sonnet 4.5 evaluated with extended thinking enabled for harder reasoning and agentic tasks. Provider: Anthropic. Released 2025-09-29. Access: API.
Unified ELO 1688 ± 8, rank #297 of 1776 rated models, from 435 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Croatian Common Sense Reasoning | 82.86 | Common Sense Reasoning Average Score (%) | 100 |
| EuroEval Czech Common Sense Reasoning | 85.14 | Common Sense Reasoning Average Score (%) | 100 |
| EuroEval Danish NLU - Angry Tweets | 64.26 | Sentiment classification Score (%) | 100 |
| EuroEval Hungarian Knowledge | 85.53 | Knowledge Average Score (%) | 100 |
| EuroEval Icelandic NLU - MIM-GOLD NER | 86.91 | Named entity recognition Score (%) | 100 |
| EuroEval Latvian Knowledge | 87.47 | Knowledge Average Score (%) | 100 |
| EuroEval Lithuanian Common Sense Reasoning | 85.65 | Common Sense Reasoning Average Score (%) | 100 |
| EuroEval Polish Common Sense Reasoning | 80.96 | Common Sense Reasoning Average Score (%) | 100 |
| EuroEval Romanian Common Sense Reasoning | 80.3 | Common Sense Reasoning Average Score (%) | 100 |
| EuroEval Serbian Common Sense Reasoning | 78.7 | Common Sense Reasoning Average Score (%) | 100 |
| EuroEval Serbian Knowledge | 91.4 | Knowledge Average Score (%) | 100 |
| EuroEval Slovak Common Sense Reasoning | 77.5 | Common Sense Reasoning Average Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-5-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.