Claude Sonnet 4.6: benchmark results
Anthropic Sonnet-tier Claude model for balanced reasoning, coding, and writing workloads. Provider: Anthropic. Released 2026-02-17. Access: API.
Unified ELO 1670 ± 1, rank #97 of 1392 rated models, from 1106 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AgentTrap | 33 | Attack success (self-reported) | 100 |
| An Empirical Study of Proactive Coding Assista | 13.57 | Pass@1 (self-reported) | 100 |
| BenchTable | 86.1 | Total Score (%) | 100 |
| BusinessCaseBench | 88.4 | Score (%) | 100 |
| CLAW-Eval | 81.4 | Score (%) | 100 |
| ChaosBench-Logic v2 | 60.1 | MCC (self-reported) | 100 |
| CodeClinic | 53.1 | Overall (self-reported) | 100 |
| EnterpriseClawBench (Claude Code) | 64.41 | Primary score (%) | 100 |
| EnterpriseClawBench (DeepAgents) | 63.18 | Primary score (%) | 100 |
| EnterpriseClawBench (OpenClaw) | 62.3 | Primary score (%) | 100 |
| EuroEval English NLU - ScaLA EN | 76.44 | Linguistic acceptability Score (%) | 100 |
| EuroEval Norwegian Common Sense Reasoning | 92.76 | Common Sense Reasoning Average Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6 · How It Works · Data refreshed daily, snapshot 2026-09-05.