Claude Sonnet 4.6 (High) — benchmark results
Claude Sonnet 4.6 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2026-02-17. Access: API.
Unified ELO 1821 ± 20, rank #125 of 1776 rated models, from 41 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| DeepResearchBench | 54.87 | Average Score | 91.4 |
| APEX v1 Medicine (MD) | 66.7 | Score (%) | 88.9 |
| Multi-turn Debate (Lechmazur) | 1597.2 | Bradley-Terry Rating | 82.5 |
| OpenCompass LLM - Reasoning | 60.6 | Score (%) | 81.8 |
| OpenCompass Reasoning - Academic | 45.6 | Score (%) | 81.8 |
| ARC-AGI-2 | 60.42 | Accuracy (%) | 81 |
| Epoch AI - ECI | 153.18 | ECI Score | 78.7 |
| ARC-AGI-1 | 86.5 | Accuracy (%) | 78.3 |
| CocoaBench | 34 | Accuracy | 77.8 |
| OpenCompass Knowledge - Engineering | 94.2 | Score (%) | 77.3 |
| OpenCompass Knowledge - Social Science | 92.5 | Score (%) | 77.3 |
| C4 Benchmark | 19.3 | Overall Score (%) | 75 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.