Claude 3.5 Sonnet: benchmark results

Anthropic Claude 3.5 Sonnet model row for later Sonnet 3.5 results. Provider: Anthropic. Released 2024-06-20. Access: API.

Unified ELO 1598 ± 1, rank #272 of 1392 rated models, from 607 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AgentLeak55.2Total Leak (self-reported)100
Arabic IFEval75.9Arabic Accuracy (%)100
CRMArena - TCU82.3TCU Score (%)100
Deception Effectiveness (Lechmazur)1.1Deception Score100
EmbodiedBench Habitat68Avg Score (%)100
Factorio Learning Environment293206Production Score100
Kazakh LLM - Constitution86.23Accuracy (%)100
Kazakh LLM - Geography64.32Accuracy (%)100
Kazakh LLM - Human Society & Rights80.27Accuracy (%)100
LLM Stats (AI2D)94.7Score (%)100
LLM Stats (ChartQA)90.8Score (%)100
MRR-Benchmark463Total Column Score100

Interactive version: theaggregate.ai/model?slug=claude-3-5-sonnet · How It Works · Data refreshed daily, snapshot 2026-09-05.