Claude Sonnet 4.6: benchmark results

Anthropic Sonnet-tier Claude model for balanced reasoning, coding, and writing workloads. Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1670 ± 1, rank #97 of 1392 rated models, from 1106 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AgentTrap33Attack success (self-reported)100
An Empirical Study of Proactive Coding Assista13.57Pass@1 (self-reported)100
BenchTable86.1Total Score (%)100
BusinessCaseBench88.4Score (%)100
CLAW-Eval81.4Score (%)100
ChaosBench-Logic v260.1MCC (self-reported)100
CodeClinic53.1Overall (self-reported)100
EnterpriseClawBench (Claude Code)64.41Primary score (%)100
EnterpriseClawBench (DeepAgents)63.18Primary score (%)100
EnterpriseClawBench (OpenClaw)62.3Primary score (%)100
EuroEval English NLU - ScaLA EN76.44Linguistic acceptability Score (%)100
EuroEval Norwegian Common Sense Reasoning92.76Common Sense Reasoning Average Score (%)100

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6 · How It Works · Data refreshed daily, snapshot 2026-09-05.