Claude Opus 5 (Claude Code): benchmark results

Provider: Anthropic. Access: API.

Unified ELO 1927 ± 20, rank #17 of 1605 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BioSecBench-Function50.3Endpoint pass rate (%)100
CliniCARE-Bench - Process Rubric Score82Process rubric score (%; GPT-5.5 judge)100
CliniCARE-Bench - Verdict Macro-F10.7Macro-F1 over the four verdict classes (0-1)100
FormalTCS - Autoformalization11.2BEq+ equivalence rate (%; natural-language theorem translate100
FormalTCS - Proof Elicitation68.7LLM-rubric score (0-100) from a Qwen3.8-Max judge in QoderCL100
FormalTCS - Theorem Proving28.7Pass@8 (%; at least one of 8 sampled Lean 4 proofs of the hu100
VariantBench49.72Endpoint pass rate (%)100
CliniCARE-Bench - Defect-Free Accuracy70Correct verdicts whose report has no judged defect (%)93.3
CliniCARE-Bench - Expected Calibration Error0.06Expected calibration error (0-1; lower is better)93.3
CliniCARE-Bench - Verdict Accuracy75.6Four-class verdict accuracy (%)93.3
InfraBench - Attempt Pass@0.586.1Attempts scoring at least 0.5 (%)88.2
InfraBench - Attempt Pass@172.2Attempts with a perfect score (%)88.2

Interactive version: theaggregate.ai/model?slug=claude-opus-5-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.