Claude Sonnet 4.6 (Claude Code): benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1737 ± 14, rank #166 of 1605 rated models, from 50 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EnvTrustBench55.3Environmental misgrounding rate (%; share of accepted pass-o100
FaulT-Bench - Wrong-Cause Tickets - Fix Score0.96Judged remediation score (0-1; answered runs)100
FaulT-Bench - Wrong-Cause Tickets - Outcome Score0.98Judged outcome score (0-1; 0.7 diagnosis accuracy + 0.3 expl100
FaulT-Bench - Wrong-Device Tickets - Fix Score0.98Judged remediation score (0-1; answered runs)100
FaulT-Bench - Wrong-Device Tickets - Outcome Score0.99Judged outcome score (0-1; 0.7 diagnosis accuracy + 0.3 expl100
FaulT-Bench - Wrong-Device Tickets - Reasoning Score0.96Judged reasoning score (0-1; grounding, causality and covera100
Herculean - Auditing (Claude Code)66.15Accuracy (%; share of the 65 SEC XBRL audit instances where 100
MDGym - Medium7.3Full success rate (%; share of the 55 medium expert-curated 100
ProcCtrlBench0.74ProcCtrlBench score PB (0-1; calibrated process quality over100
SREGym (No Noise)60.7End-to-end success rate (%; runs with both a correct diagnos100
SREGym (No Noise) - Diagnosis72.6Diagnosis success rate (%; checklist-based LLM-judge verdict100
SREGym (Noise Injected)53.7End-to-end success rate (%; runs with both a correct diagnos100

Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.