Claude Opus 4.5: benchmark results

Anthropic Opus-tier Claude model for complex reasoning, coding, and writing. Provider: Anthropic. Released 2025-11-24. Access: API.

Unified ELO 1685 ± 1, rank #76 of 1392 rated models, from 344 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AGC-Bench - scar1.12Dataset z-score100
AGC-Bench - speak_to_structure1.51Dataset z-score100
ASCIIBench1702ELO Rating100
ATM-Bench86Oracle QS (%)100
Cited but Not Verified95.7Relevant Content (self-reported)100
FormalRewardBench59.8Pairwise Accuracy (self-reported)100
HAL CORE-Bench Hard77.78Accuracy (%)100
MMLongBench-Doc - Accuracy61.9Accuracy (%)100
MonitoringBench95.1Baseline catch rate (%) (self-reported)100
OpenClaw Arena Model Leaderboard67.4Avg Score (self-reported)100
Phare - Hallucination Resistance88.23Score (%)100
SecCodeBench68.1Total Score100

Interactive version: theaggregate.ai/model?slug=claude-opus-4-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.