Claude Opus 4.5 (Thinking 32K): benchmark results
Provider: Anthropic. Released 2025-11-24. Access: API.
Unified ELO 1694 ± 1, rank #159 of 3078 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Korean CSAT 2026 (Easy Mode) - Korean | 100 | Points (out of 100) | 92.5 |
| Korean CSAT 2026 (Easy Mode) - Mathematics | 100 | Points (out of 100) | 88.4 |
| Korean CSAT 2026 (Easy Mode) - Life Science I | 48 | Points (out of 50) | 80.3 |
| Korean CSAT 2026 (Easy Mode) - Total | 429.5 | Points (out of 450) | 76.2 |
| Korean CSAT 2026 (Easy Mode) - English | 97 | Points (out of 100) | 66 |
| Korean CSAT 2026 (Easy Mode) - Chemistry I | 45 | Points (out of 50) | 63.3 |
| FrontierMath - Tiers 1-3 | 20.69 | Accuracy (%, 290 problems) | 61.1 |
| Korean CSAT 2026 (Easy Mode) - Physics I | 33 | Points (out of 50) | 59.2 |
| ARC-AGI-1 | 75.83 | Accuracy (%) | 57.3 |
| Korean CSAT 2026 (Easy Mode) - Society and Culture | 39 | Points (out of 50) | 55.8 |
| FrontierMath - Tiers 1-3 (v2) | 34.39 | Accuracy (%, 285 private v2 problems) | 53.1 |
| FrontierMath - Tier 4 | 4.17 | Accuracy (%, 48 problems) | 44.4 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-5-thinking-32k · How It Works · Data refreshed daily, snapshot 2026-09-19.