Claude Opus 4.5 (20251101) (Thinking): benchmark results

Claude Opus 4.5 (20251101) evaluated with thinking enabled. Provider: Anthropic. Released 2025-11-01. Access: API.

Unified ELO 1671 ± 1, rank #175 of 1761 rated models, from 22 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Vals AI MGSM95.2Accuracy (%)100
UGI - Writing70.34Writing Score98.5
UGI - Natural Intelligence65.22NatInt Score96
SEAL - MASK92.53Score93.9
Vals AI MedQA95.88Accuracy (%)90.4
Vals AI AIME95.42Accuracy (%)88.4
SEAL - Humanity's Last Exam (Text Only)26.32Score84.2
SEAL - Humanity's Last Exam25.2Score78
EnigmaEval11.91Score (self-reported)77
SEAL - MultiNRC48.63Score74.4
Vals AI CaseLaw v262.59Accuracy (%)69.5
PM-LLM-Benchmark33.2Score67.8

Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-20251101-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.