Claude Opus 4.5 — benchmark results

Anthropic Opus-tier Claude model for complex reasoning, coding, and writing. Provider: Anthropic. Released 2025-11-24. Access: API.

Unified ELO 1800 ± 13, rank #140 of 1776 rated models, from 283 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AGC-Bench - scar1.12Dataset z-score100
AGC-Bench - speak_to_structure1.51Dataset z-score100
Cited but Not Verified95.7Relevant Content (self-reported)100
FormalRewardBench59.8Pairwise Accuracy (self-reported)100
HAL CORE-Bench Hard77.78Accuracy (%)100
MMLongBench-Doc - Accuracy61.9Accuracy (%)100
MonitoringBench95.1Baseline catch rate (%) (self-reported)100
OpenClaw Arena Model Leaderboard67.4Avg Score (self-reported)100
SecCodeBench68.1Total Score100
The Metanym Game7T (self-reported)100
Vals AI MGSM95.2Accuracy (%)99.6
AGC-Bench - metaphor_generation2.4Dataset z-score98.8

Interactive version: theaggregate.ai/model?slug=claude-opus-4-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.