Claude Opus 4.5 (High) — benchmark results
Claude Opus 4.5 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2025-11-24. Access: API.
Unified ELO 1732 ± 28, rank #219 of 1776 rated models, from 40 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| NonoBench | 56.7 | Overall Accuracy (%) | 93.6 |
| MedScribe | 85.32 | Score (self-reported) | 86.5 |
| MCPMark | 42.32 | Pass@1 (%) | 84.2 |
| Finance Agent v1.1 | 58.81 | Score (self-reported) | 78.2 |
| MedCode | 49.16 | Score (self-reported) | 77.4 |
| Chess Bench LLM | 1124 | Lichess Rating | 75.8 |
| HAL CORE-Bench Hard | 42.22 | Accuracy (%) | 75 |
| IOI | 23.58 | Score (self-reported) | 69.8 |
| Vals AI ProofBench | 36 | Accuracy (%) | 69.7 |
| APEX-Agents | 34.8 | Mean Score (ReAct) (self-reported) | 62.8 |
| VoxelBench | 1477 | Rating | 61.7 |
| APEX v1 Medicine (MD) | 61.5 | Score (%) | 61.1 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.