Claude Opus 4.5 (Thinking) — benchmark results
Claude Opus 4.5 evaluated with thinking enabled. Provider: Anthropic. Released 2025-11-24. Access: API.
Unified ELO 1847 ± 15, rank #106 of 1776 rated models, from 120 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MCP-Atlas (Opus 4.6 System Card) | 62.3 | Score (%) | 100 |
| OpenCompass Agent - Multi-Turn | 84.8 | Score (%) | 100 |
| SWE-bench Verified (Opus 4.6 System Card) | 80.9 | Resolved (%) | 100 |
| AA MMLU-Pro | 89.45 | Accuracy (%) | 99.4 |
| AA Long Context Reasoning | 74 | Accuracy (%) | 98.3 |
| AA LiveCodeBench | 87.09 | Pass@1 (%) | 98 |
| Wolfram LLM Benchmarking Project | 68.9 | Correct Functionality (%) | 98 |
| AA Global-MMLU-Lite - German | 93 | Accuracy (%) | 97.5 |
| AA Global-MMLU-Lite - Burmese | 89.17 | Accuracy (%) | 97.3 |
| AA Omniscience - Software Engineering (SWE) - Julia | 72 | Accuracy (%) | 97.1 |
| AA Global-MMLU-Lite - Korean | 91.58 | Accuracy (%) | 97 |
| Pencil Puzzle Bench - Hitori | 40 | Direct-ask Success Rate (%) | 97 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.