Claude Opus 4.5 (20251101) (Thinking) — benchmark results
Claude Opus 4.5 (20251101) evaluated with thinking enabled. Provider: Anthropic. Released 2025-11-01. Access: API.
Unified ELO 1805 ± 14, rank #138 of 1776 rated models, from 38 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Vals AI MGSM | 95.2 | Accuracy (%) | 99.6 |
| UGI - Writing | 70.34 | Writing Score | 98.9 |
| UGI - Natural Intelligence | 65.22 | NatInt Score | 96.6 |
| SEAL - MASK | 92.53 | Score | 93.9 |
| Vals AI SAGE | 52.09 | Accuracy (%) | 92.6 |
| Vals AI MedQA | 95.88 | Accuracy (%) | 90.4 |
| Vals AI AIME | 95.42 | Accuracy (%) | 88.4 |
| SEAL - EnigmaEval | 11.91 | Score | 87.5 |
| Vals AI Finance Agent | 58.81 | Accuracy (%) | 86 |
| Vals AI LegalBench | 84.6 | Accuracy (%) | 84.8 |
| Vals AI TaxEval v2 | 74.86 | Accuracy (%) | 84.8 |
| Vals AI MedScribe | 85.32 | Accuracy (%) | 83.3 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-20251101-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.