Claude Sonnet 4.6 (High): benchmark results

Claude Sonnet 4.6 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1669 ± 1, rank #182 of 1761 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
DeepResearchBench54.9Average Score97.5
CocoaBench34Accuracy77.8
Multi-turn Debate (Lechmazur)1585.2Bradley-Terry Rating76.7
APEX-Agents40.7Mean Score (ReAct) (self-reported)75
C4 Benchmark19.3Overall Score (%)75
ARC-AGI-260.42Accuracy (%)74.8
GRIPS92.6Accuracy (%)73.1
GIM1.12IRT ability (theta)71.1
ARC-AGI-186.5Accuracy (%)69.5
GRIPS-hard60.4Accuracy (%)69.2
PACT (Lechmazur)1546PACT Bilateral Rating64
OTIS Mock AIME 2024-2575.56Accuracy (%)59.9

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-high · How It Works · Data refreshed daily, snapshot 2026-09-05.