Claude Opus 4.1 (Thinking) — benchmark results
Claude Opus 4.1 evaluated with thinking enabled. Provider: Anthropic. Released 2025-08-05. Access: API.
Unified ELO 1746 ± 44, rank #203 of 1776 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenchTable | 82.7 | Total Score (%) | 98.8 |
| AA MMLU-Pro | 87.99 | Accuracy (%) | 98.3 |
| Wolfram LLM Benchmarking Project | 64.7 | Correct Functionality (%) | 94.5 |
| Ducky Bench (Stabby Quack) | 1399 | ELO | 89.3 |
| AA Long Context Reasoning | 66.33 | Accuracy (%) | 85.4 |
| Artificial Analysis Intelligence Index | 33.71 | Intelligence Index | 81.5 |
| AA Terminal-Bench Hard | 34.34 | Accuracy (%) | 79.4 |
| MathVision | 66 | Overall Accuracy (%) | 79.2 |
| AA SciCode | 40.9 | Accuracy (%) | 77.3 |
| AA AIME 2025 | 80.33 | Accuracy (%) | 75.9 |
| ZeroBench | 5 | Score (%) | 75.4 |
| AA GPQA Diamond | 80.9 | Accuracy (%) | 74.1 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.