Claude Opus 4.5 (High): benchmark results
Claude Opus 4.5 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2025-11-24. Access: API.
Unified ELO 1639 ± 1, rank #286 of 1761 rated models, from 32 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SWE-bench Verified | 76.8 | Resolved (%) | 100 |
| NonoBench | 56.7 | Overall Accuracy (%) | 93.6 |
| MCPMark | 42.32 | Pass@1 (%) | 84.2 |
| Chess Bench LLM | 1009 | Lichess Rating | 78.3 |
| HAL CORE-Bench Hard | 42.22 | Accuracy (%) | 75 |
| SlopCodeBench | 17.35 | Isolated Solved (%) | 56.2 |
| APEX-Agents | 34.8 | Mean Score (ReAct) (self-reported) | 55.7 |
| ATLAS | 51.85 | ATLAScore (self-reported) | 54.5 |
| ChartMuseum | 60.7 | Overall Accuracy (%) | 52.4 |
| Pencil Puzzle Bench - Heyawake | 0 | Direct-ask Success Rate (%) | 50 |
| Pencil Puzzle Bench - Sashigane | 0 | Direct-ask Success Rate (%) | 50 |
| Pencil Puzzle Bench - Shakashaka | 0 | Direct-ask Success Rate (%) | 50 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-high · How It Works · Data refreshed daily, snapshot 2026-09-05.