Claude 3.5 Sonnet — benchmark results

Anthropic Claude 3.5 Sonnet model row for later Sonnet 3.5 results. Provider: Anthropic. Released 2024-06-20. Access: API.

Unified ELO 1533 ± 9, rank #670 of 1776 rated models, from 591 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Aider Refactoring Benchmark92.1Percent completed correctly (self-reported)100
Arabic IFEval75.9Arabic Accuracy (%)100
CRMArena - TCU82.3TCU Score (%)100
Clembench Multimodal v1.6.580.77clemscore (self-reported)100
ComplexFuncBench61Score (self-reported)100
Deception Effectiveness (Lechmazur)1.1Deception Score100
EmbodiedBench Habitat68Avg Score (%)100
Factorio Learning Environment293206Production Score100
Kazakh LLM - Constitution86.23Accuracy (%)100
Kazakh LLM - Geography64.32Accuracy (%)100
Kazakh LLM - Human Society & Rights80.27Accuracy (%)100
LLM Stats (AI2D)94.7Score (%)100

Interactive version: theaggregate.ai/model?slug=claude-3-5-sonnet · How the rankings work · Data refreshed daily, snapshot 2026-07-22.