Claude Sonnet 5 (Claude Code): benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1876 ± 20, rank #36 of 1605 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CliniCARE-Bench - Expected Calibration Error | 0.05 | Expected calibration error (0-1; lower is better) | 100 |
| CliniCARE-Bench - Over-Commitment Rate | 30.7 | Definitive verdicts on cases whose reference requires deferr | 93.3 |
| CliniCARE-Bench - Verdict Accuracy | 73.9 | Four-class verdict accuracy (%) | 86.7 |
| CliniCARE-Bench - Defect-Free Accuracy | 65.6 | Correct verdicts whose report has no judged defect (%) | 76.7 |
| InfraBench - Attempt Pass@0.5 | 80.6 | Attempts scoring at least 0.5 (%) | 76.5 |
| CliniCARE-Bench - Verdict Macro-F1 | 0.65 | Macro-F1 over the four verdict classes (0-1) | 73.3 |
| VariantBench | 38.7 | Endpoint pass rate (%) | 71.7 |
| FormalTCS - Autoformalization | 9.1 | BEq+ equivalence rate (%; natural-language theorem translate | 71.4 |
| FormalTCS - Proof Elicitation | 65.7 | LLM-rubric score (0-100) from a Qwen3.8-Max judge in QoderCL | 71.4 |
| FormalTCS - Theorem Elicitation | 63 | LLM-rubric score (0-100) from a Qwen3.8-Max judge in QoderCL | 71.4 |
| FormalTCS - Theorem Proving | 23.8 | Pass@8 (%; at least one of 8 sampled Lean 4 proofs of the hu | 71.4 |
| InfraBench - Best-of-3 Perfect Tasks | 75 | Tasks with a perfect best-of-three attempt (%; of 12) | 67.6 |
Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.