Claude Sonnet 4.6 (High): benchmark results
Claude Sonnet 4.6 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2026-02-17. Access: API.
Unified ELO 1669 ± 1, rank #182 of 1761 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| DeepResearchBench | 54.9 | Average Score | 97.5 |
| CocoaBench | 34 | Accuracy | 77.8 |
| Multi-turn Debate (Lechmazur) | 1585.2 | Bradley-Terry Rating | 76.7 |
| APEX-Agents | 40.7 | Mean Score (ReAct) (self-reported) | 75 |
| C4 Benchmark | 19.3 | Overall Score (%) | 75 |
| ARC-AGI-2 | 60.42 | Accuracy (%) | 74.8 |
| GRIPS | 92.6 | Accuracy (%) | 73.1 |
| GIM | 1.12 | IRT ability (theta) | 71.1 |
| ARC-AGI-1 | 86.5 | Accuracy (%) | 69.5 |
| GRIPS-hard | 60.4 | Accuracy (%) | 69.2 |
| PACT (Lechmazur) | 1546 | PACT Bilateral Rating | 64 |
| OTIS Mock AIME 2024-25 | 75.56 | Accuracy (%) | 59.9 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-high · How It Works · Data refreshed daily, snapshot 2026-09-05.