Claude Opus 4.5 (High): benchmark results

Claude Opus 4.5 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2025-11-24. Access: API.

Unified ELO 1639 ± 1, rank #286 of 1761 rated models, from 32 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SWE-bench Verified76.8Resolved (%)100
NonoBench56.7Overall Accuracy (%)93.6
MCPMark42.32Pass@1 (%)84.2
Chess Bench LLM1009Lichess Rating78.3
HAL CORE-Bench Hard42.22Accuracy (%)75
SlopCodeBench17.35Isolated Solved (%)56.2
APEX-Agents34.8Mean Score (ReAct) (self-reported)55.7
ATLAS51.85ATLAScore (self-reported)54.5
ChartMuseum60.7Overall Accuracy (%)52.4
Pencil Puzzle Bench - Heyawake0Direct-ask Success Rate (%)50
Pencil Puzzle Bench - Sashigane0Direct-ask Success Rate (%)50
Pencil Puzzle Bench - Shakashaka0Direct-ask Success Rate (%)50

Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-high · How It Works · Data refreshed daily, snapshot 2026-09-05.