Claude Opus 4.1 (20250805) (Thinking): benchmark results
Claude Opus 4.1 (20250805) evaluated with thinking enabled. Provider: Anthropic. Released 2025-08-05. Access: API.
Unified ELO 1631 ± 1, rank #312 of 1761 rated models, from 36 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Vals AI MGSM | 94.44 | Accuracy (%) | 97.7 |
| UGI - Writing | 66.51 | Writing Score | 96.4 |
| Vals AI MATH 500 | 95.4 | Accuracy (%) | 95.8 |
| SEAL - MASK | 94.2 | Score | 95.5 |
| UGI - Natural Intelligence | 58.17 | NatInt Score | 93 |
| Vals AI MMLU-Pro | 87.92 | Accuracy (%) | 83.9 |
| Vals AI MedQA | 93.59 | Accuracy (%) | 77.7 |
| SEAL - VISTA | 48.44 | Score | 74.2 |
| MedCode | 47.23 | Score (self-reported) | 71.6 |
| Vals AI MedCode | 47.23 | Accuracy (%) | 69.7 |
| SEAL Showdown | 1093 | Arena Score | 69.6 |
| Vals AI TaxEval v2 | 73.67 | Accuracy (%) | 68.1 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-20250805-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.