Claude 3.5 Sonnet: benchmark results
Anthropic Claude 3.5 Sonnet model row for later Sonnet 3.5 results. Provider: Anthropic. Released 2024-06-20. Access: API.
Unified ELO 1598 ± 1, rank #272 of 1392 rated models, from 607 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AgentLeak | 55.2 | Total Leak (self-reported) | 100 |
| Arabic IFEval | 75.9 | Arabic Accuracy (%) | 100 |
| CRMArena - TCU | 82.3 | TCU Score (%) | 100 |
| Deception Effectiveness (Lechmazur) | 1.1 | Deception Score | 100 |
| EmbodiedBench Habitat | 68 | Avg Score (%) | 100 |
| Factorio Learning Environment | 293206 | Production Score | 100 |
| Kazakh LLM - Constitution | 86.23 | Accuracy (%) | 100 |
| Kazakh LLM - Geography | 64.32 | Accuracy (%) | 100 |
| Kazakh LLM - Human Society & Rights | 80.27 | Accuracy (%) | 100 |
| LLM Stats (AI2D) | 94.7 | Score (%) | 100 |
| LLM Stats (ChartQA) | 90.8 | Score (%) | 100 |
| MRR-Benchmark | 463 | Total Column Score | 100 |
Interactive version: theaggregate.ai/model?slug=claude-3-5-sonnet · How It Works · Data refreshed daily, snapshot 2026-09-05.