Claude Opus 4.5 — benchmark results
Anthropic Opus-tier Claude model for complex reasoning, coding, and writing. Provider: Anthropic. Released 2025-11-24. Access: API.
Unified ELO 1800 ± 13, rank #140 of 1776 rated models, from 283 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - scar | 1.12 | Dataset z-score | 100 |
| AGC-Bench - speak_to_structure | 1.51 | Dataset z-score | 100 |
| Cited but Not Verified | 95.7 | Relevant Content (self-reported) | 100 |
| FormalRewardBench | 59.8 | Pairwise Accuracy (self-reported) | 100 |
| HAL CORE-Bench Hard | 77.78 | Accuracy (%) | 100 |
| MMLongBench-Doc - Accuracy | 61.9 | Accuracy (%) | 100 |
| MonitoringBench | 95.1 | Baseline catch rate (%) (self-reported) | 100 |
| OpenClaw Arena Model Leaderboard | 67.4 | Avg Score (self-reported) | 100 |
| SecCodeBench | 68.1 | Total Score | 100 |
| The Metanym Game | 7 | T (self-reported) | 100 |
| Vals AI MGSM | 95.2 | Accuracy (%) | 99.6 |
| AGC-Bench - metaphor_generation | 2.4 | Dataset z-score | 98.8 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.