Claude Opus 4.5 (Thinking) — benchmark results

Claude Opus 4.5 evaluated with thinking enabled. Provider: Anthropic. Released 2025-11-24. Access: API.

Unified ELO 1847 ± 15, rank #106 of 1776 rated models, from 120 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MCP-Atlas (Opus 4.6 System Card)62.3Score (%)100
OpenCompass Agent - Multi-Turn84.8Score (%)100
SWE-bench Verified (Opus 4.6 System Card)80.9Resolved (%)100
AA MMLU-Pro89.45Accuracy (%)99.4
AA Long Context Reasoning74Accuracy (%)98.3
AA LiveCodeBench87.09Pass@1 (%)98
Wolfram LLM Benchmarking Project68.9Correct Functionality (%)98
AA Global-MMLU-Lite - German93Accuracy (%)97.5
AA Global-MMLU-Lite - Burmese89.17Accuracy (%)97.3
AA Omniscience - Software Engineering (SWE) - Julia72Accuracy (%)97.1
AA Global-MMLU-Lite - Korean91.58Accuracy (%)97
Pencil Puzzle Bench - Hitori40Direct-ask Success Rate (%)97

Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.