Claude Opus 4.1 (20250805) — benchmark results
August 5, 2025 Claude Opus 4.1 snapshot, tracked when sources report the dated API model. Provider: Anthropic. Released 2025-08-05. Access: API.
Unified ELO 1723 ± 10, rank #233 of 1776 rated models, from 60 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| WebApp1K Duo | 76.9 | Pass@1 (%) | 95.8 |
| Vals AI MGSM | 94.22 | Accuracy (%) | 94 |
| HAL CORE-Bench Hard | 51.11 | Accuracy (%) | 92.3 |
| HAL SWE-bench Verified Mini | 61 | Score (%) | 88.2 |
| HAL SciCode | 7.69 | Accuracy (%) | 86.7 |
| SEAL - MASK | 87.4 | Score | 84.8 |
| Chatbot Arena (Text) | 1447 | Elo | 84.7 |
| HAL TAU-bench Airline | 54 | Accuracy (%) | 82.1 |
| Arabic Broad Leaderboard | 8.88 | Average Score (0-10) | 81.9 |
| Vals AI MMLU-Pro | 87.21 | Accuracy (%) | 80.2 |
| ForecastBench | 65.2 | Overall Score (higher is better) | 78.7 |
| HAL GAIA Level 2 | 66.28 | Accuracy (%) | 75 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-20250805 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.