Claude Sonnet 4.6 (High) — benchmark results

Claude Sonnet 4.6 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2026-02-17. Access: API.

Unified ELO 1821 ± 20, rank #125 of 1776 rated models, from 41 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
DeepResearchBench54.87Average Score91.4
APEX v1 Medicine (MD)66.7Score (%)88.9
Multi-turn Debate (Lechmazur)1597.2Bradley-Terry Rating82.5
OpenCompass LLM - Reasoning60.6Score (%)81.8
OpenCompass Reasoning - Academic45.6Score (%)81.8
ARC-AGI-260.42Accuracy (%)81
Epoch AI - ECI153.18ECI Score78.7
ARC-AGI-186.5Accuracy (%)78.3
CocoaBench34Accuracy77.8
OpenCompass Knowledge - Engineering94.2Score (%)77.3
OpenCompass Knowledge - Social Science92.5Score (%)77.3
C4 Benchmark19.3Overall Score (%)75

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.