Claude Opus 4.1 (High) — benchmark results
Claude Opus 4.1 evaluated at the high reasoning-effort setting. Provider: Anthropic. Released 2025-08-05. Access: API.
Unified ELO 1787 ± 13, rank #154 of 1776 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HAL GAIA | 68.48 | Accuracy (%) | 87.5 |
| HAL GAIA Level 2 | 70.93 | Accuracy (%) | 87.5 |
| HAL SciCode | 6.92 | Accuracy (%) | 80 |
| HAL GAIA Level 3 | 53.85 | Accuracy (%) | 78.1 |
| HAL SWE-bench Verified Mini | 54 | Score (%) | 76.5 |
| HAL CORE-Bench Hard | 42.22 | Accuracy (%) | 75 |
| HAL USACO | 51.47 | Accuracy (%) | 72.7 |
| HAL GAIA Level 1 | 71.7 | Accuracy (%) | 71.9 |
| HAL TAU-bench Airline | 52 | Accuracy (%) | 67.9 |
| HAL AssistantBench | 13.75 | Accuracy (%) | 57.1 |
| HAL ScienceAgentBench | 26.47 | Accuracy (%) | 46.7 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.