Claude Sonnet 4.6 (Claude Code): benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1737 ± 14, rank #166 of 1605 rated models, from 50 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EnvTrustBench | 55.3 | Environmental misgrounding rate (%; share of accepted pass-o | 100 |
| FaulT-Bench - Wrong-Cause Tickets - Fix Score | 0.96 | Judged remediation score (0-1; answered runs) | 100 |
| FaulT-Bench - Wrong-Cause Tickets - Outcome Score | 0.98 | Judged outcome score (0-1; 0.7 diagnosis accuracy + 0.3 expl | 100 |
| FaulT-Bench - Wrong-Device Tickets - Fix Score | 0.98 | Judged remediation score (0-1; answered runs) | 100 |
| FaulT-Bench - Wrong-Device Tickets - Outcome Score | 0.99 | Judged outcome score (0-1; 0.7 diagnosis accuracy + 0.3 expl | 100 |
| FaulT-Bench - Wrong-Device Tickets - Reasoning Score | 0.96 | Judged reasoning score (0-1; grounding, causality and covera | 100 |
| Herculean - Auditing (Claude Code) | 66.15 | Accuracy (%; share of the 65 SEC XBRL audit instances where | 100 |
| MDGym - Medium | 7.3 | Full success rate (%; share of the 55 medium expert-curated | 100 |
| ProcCtrlBench | 0.74 | ProcCtrlBench score PB (0-1; calibrated process quality over | 100 |
| SREGym (No Noise) | 60.7 | End-to-end success rate (%; runs with both a correct diagnos | 100 |
| SREGym (No Noise) - Diagnosis | 72.6 | Diagnosis success rate (%; checklist-based LLM-judge verdict | 100 |
| SREGym (Noise Injected) | 53.7 | End-to-end success rate (%; runs with both a correct diagnos | 100 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-4-6-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.