Claude 3.5 Sonnet — benchmark results
Anthropic Claude 3.5 Sonnet model row for later Sonnet 3.5 results. Provider: Anthropic. Released 2024-06-20. Access: API.
Unified ELO 1533 ± 9, rank #670 of 1776 rated models, from 591 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Aider Refactoring Benchmark | 92.1 | Percent completed correctly (self-reported) | 100 |
| Arabic IFEval | 75.9 | Arabic Accuracy (%) | 100 |
| CRMArena - TCU | 82.3 | TCU Score (%) | 100 |
| Clembench Multimodal v1.6.5 | 80.77 | clemscore (self-reported) | 100 |
| ComplexFuncBench | 61 | Score (self-reported) | 100 |
| Deception Effectiveness (Lechmazur) | 1.1 | Deception Score | 100 |
| EmbodiedBench Habitat | 68 | Avg Score (%) | 100 |
| Factorio Learning Environment | 293206 | Production Score | 100 |
| Kazakh LLM - Constitution | 86.23 | Accuracy (%) | 100 |
| Kazakh LLM - Geography | 64.32 | Accuracy (%) | 100 |
| Kazakh LLM - Human Society & Rights | 80.27 | Accuracy (%) | 100 |
| LLM Stats (AI2D) | 94.7 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-3-5-sonnet · How the rankings work · Data refreshed daily, snapshot 2026-07-22.