Claude Opus 4.5: benchmark results
Anthropic Opus-tier Claude model for complex reasoning, coding, and writing. Provider: Anthropic. Released 2025-11-24. Access: API.
Unified ELO 1685 ± 1, rank #76 of 1392 rated models, from 344 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - scar | 1.12 | Dataset z-score | 100 |
| AGC-Bench - speak_to_structure | 1.51 | Dataset z-score | 100 |
| ASCIIBench | 1702 | ELO Rating | 100 |
| ATM-Bench | 86 | Oracle QS (%) | 100 |
| Cited but Not Verified | 95.7 | Relevant Content (self-reported) | 100 |
| FormalRewardBench | 59.8 | Pairwise Accuracy (self-reported) | 100 |
| HAL CORE-Bench Hard | 77.78 | Accuracy (%) | 100 |
| MMLongBench-Doc - Accuracy | 61.9 | Accuracy (%) | 100 |
| MonitoringBench | 95.1 | Baseline catch rate (%) (self-reported) | 100 |
| OpenClaw Arena Model Leaderboard | 67.4 | Avg Score (self-reported) | 100 |
| Phare - Hallucination Resistance | 88.23 | Score (%) | 100 |
| SecCodeBench | 68.1 | Total Score | 100 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.