Claude Opus 4.1 (20250805) (Thinking) — benchmark results
Claude Opus 4.1 (20250805) evaluated with thinking enabled. Provider: Anthropic. Released 2025-08-05. Access: API.
Unified ELO 1744 ± 17, rank #207 of 1776 rated models, from 34 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Writing | 66.51 | Writing Score | 97 |
| Vals AI MGSM | 94.44 | Accuracy (%) | 97 |
| SEAL - MASK | 94.2 | Score | 95.5 |
| Vals AI MATH 500 | 95.4 | Accuracy (%) | 94.1 |
| UGI - Natural Intelligence | 58.17 | NatInt Score | 93.7 |
| Vals AI MMLU-Pro | 87.92 | Accuracy (%) | 88.4 |
| Vals AI MedQA | 93.59 | Accuracy (%) | 77.7 |
| SEAL - VISTA | 48.44 | Score | 74.2 |
| Vals AI MedCode | 47.23 | Accuracy (%) | 72.6 |
| SEAL Showdown | 1099.6 | Arena Score | 71.7 |
| Vals AI TaxEval v2 | 73.67 | Accuracy (%) | 71.1 |
| SEAL - EnigmaEval | 7.18 | Score | 69.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-20250805-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.