Claude Opus 4.5 (20251101) (Thinking 32K): benchmark results
Claude Opus 4.5 (20251101) evaluated with a 32K-token thinking budget. Provider: Anthropic. Released 2025-11-01. Access: API.
Unified ELO 1671 ± 1, rank #177 of 1761 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| OTIS Mock AIME 2024-25 | 86.11 | Accuracy (%) | 71.9 |
| WebDev Arena | 1494.39 | Arena Score | 65 |
| SimpleQA Verified | 45.7 | Accuracy (%) | 61 |
| FrontierMath - Tiers 1-3 (v2) | 34.39 | Accuracy (%, 285 private v2 problems) | 54.3 |
| VPCT | 40 | Accuracy (%) | 50 |
| Chess Puzzles (Epoch AI) | 12 | Accuracy (%) | 43.6 |
| FrontierMath - Tier 4 (v2) | 4.88 | Accuracy (%, 41 private v2 problems) | 19.1 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-20251101-thinking-32k · How It Works · Data refreshed daily, snapshot 2026-09-05.