Claude Opus 4.5 (20251101) (Thinking) — benchmark results

Claude Opus 4.5 (20251101) evaluated with thinking enabled. Provider: Anthropic. Released 2025-11-01. Access: API.

Unified ELO 1805 ± 14, rank #138 of 1776 rated models, from 38 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Vals AI MGSM95.2Accuracy (%)99.6
UGI - Writing70.34Writing Score98.9
UGI - Natural Intelligence65.22NatInt Score96.6
SEAL - MASK92.53Score93.9
Vals AI SAGE52.09Accuracy (%)92.6
Vals AI MedQA95.88Accuracy (%)90.4
Vals AI AIME95.42Accuracy (%)88.4
SEAL - EnigmaEval11.91Score87.5
Vals AI Finance Agent58.81Accuracy (%)86
Vals AI LegalBench84.6Accuracy (%)84.8
Vals AI TaxEval v274.86Accuracy (%)84.8
Vals AI MedScribe85.32Accuracy (%)83.3

Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-20251101-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.