Claude Opus 5 (Claude Code): benchmark results
Provider: Anthropic. Access: API.
Unified ELO 1927 ± 20, rank #17 of 1605 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BioSecBench-Function | 50.3 | Endpoint pass rate (%) | 100 |
| CliniCARE-Bench - Process Rubric Score | 82 | Process rubric score (%; GPT-5.5 judge) | 100 |
| CliniCARE-Bench - Verdict Macro-F1 | 0.7 | Macro-F1 over the four verdict classes (0-1) | 100 |
| FormalTCS - Autoformalization | 11.2 | BEq+ equivalence rate (%; natural-language theorem translate | 100 |
| FormalTCS - Proof Elicitation | 68.7 | LLM-rubric score (0-100) from a Qwen3.8-Max judge in QoderCL | 100 |
| FormalTCS - Theorem Proving | 28.7 | Pass@8 (%; at least one of 8 sampled Lean 4 proofs of the hu | 100 |
| VariantBench | 49.72 | Endpoint pass rate (%) | 100 |
| CliniCARE-Bench - Defect-Free Accuracy | 70 | Correct verdicts whose report has no judged defect (%) | 93.3 |
| CliniCARE-Bench - Expected Calibration Error | 0.06 | Expected calibration error (0-1; lower is better) | 93.3 |
| CliniCARE-Bench - Verdict Accuracy | 75.6 | Four-class verdict accuracy (%) | 93.3 |
| InfraBench - Attempt Pass@0.5 | 86.1 | Attempts scoring at least 0.5 (%) | 88.2 |
| InfraBench - Attempt Pass@1 | 72.2 | Attempts with a perfect score (%) | 88.2 |
Interactive version: theaggregate.ai/model?slug=claude-opus-5-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-26.