Claude Opus 4 (Non-reasoning) — benchmark results
Claude Opus 4 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1675 ± 30, rank #311 of 1776 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Writing | 68.24 | Writing Score | 98.2 |
| Generalization V1 (Lechmazur) | 1.7 | Avg Rank (lower is better) | 95.6 |
| Confabulation Leaderboard (Lechmazur) | 2.97 | Confabulation rate % (lower is better) | 95.2 |
| AA MMLU-Pro | 85.98 | Accuracy (%) | 93.8 |
| UGI - Natural Intelligence | 57.57 | NatInt Score | 93.7 |
| AA SciCode | 40.86 | Accuracy (%) | 76.9 |
| AA MATH-500 | 94.07 | Accuracy (%) | 76.5 |
| Artificial Analysis Intelligence Index | 25.5 | Intelligence Index | 69.3 |
| NYT Connections Older Models | 34.4 | Score (%) | 62 |
| Elimination Game (Lechmazur) | 4.41 | TrueSkill μ | 61 |
| AA LiveCodeBench | 54.18 | Pass@1 (%) | 60.5 |
| AA GPQA Diamond | 70.1 | Accuracy (%) | 53.6 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.