Claude Sonnet 5 (Claude Code): benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1876 ± 20, rank #36 of 1605 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CliniCARE-Bench - Expected Calibration Error0.05Expected calibration error (0-1; lower is better)100
CliniCARE-Bench - Over-Commitment Rate30.7Definitive verdicts on cases whose reference requires deferr93.3
CliniCARE-Bench - Verdict Accuracy73.9Four-class verdict accuracy (%)86.7
CliniCARE-Bench - Defect-Free Accuracy65.6Correct verdicts whose report has no judged defect (%)76.7
InfraBench - Attempt Pass@0.580.6Attempts scoring at least 0.5 (%)76.5
CliniCARE-Bench - Verdict Macro-F10.65Macro-F1 over the four verdict classes (0-1)73.3
VariantBench38.7Endpoint pass rate (%)71.7
FormalTCS - Autoformalization9.1BEq+ equivalence rate (%; natural-language theorem translate71.4
FormalTCS - Proof Elicitation65.7LLM-rubric score (0-100) from a Qwen3.8-Max judge in QoderCL71.4
FormalTCS - Theorem Elicitation63LLM-rubric score (0-100) from a Qwen3.8-Max judge in QoderCL71.4
FormalTCS - Theorem Proving23.8Pass@8 (%; at least one of 8 sampled Lean 4 proofs of the hu71.4
InfraBench - Best-of-3 Perfect Tasks75Tasks with a perfect best-of-three attempt (%; of 12)67.6

Interactive version: theaggregate.ai/model?slug=claude-sonnet-5-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.