Claude Opus 4.5 (20251101) (Thinking): benchmark results
Claude Opus 4.5 (20251101) evaluated with thinking enabled. Provider: Anthropic. Released 2025-11-01. Access: API.
Unified ELO 1671 ± 1, rank #175 of 1761 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Vals AI MGSM | 95.2 | Accuracy (%) | 100 |
| UGI - Writing | 70.34 | Writing Score | 98.5 |
| UGI - Natural Intelligence | 65.22 | NatInt Score | 96 |
| SEAL - MASK | 92.53 | Score | 93.9 |
| Vals AI MedQA | 95.88 | Accuracy (%) | 90.4 |
| Vals AI AIME | 95.42 | Accuracy (%) | 88.4 |
| SEAL - Humanity's Last Exam (Text Only) | 26.32 | Score | 84.2 |
| SEAL - Humanity's Last Exam | 25.2 | Score | 78 |
| EnigmaEval | 11.91 | Score (self-reported) | 77 |
| SEAL - MultiNRC | 48.63 | Score | 74.4 |
| Vals AI CaseLaw v2 | 62.59 | Accuracy (%) | 69.5 |
| PM-LLM-Benchmark | 33.2 | Score | 67.8 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-20251101-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.