Claude Opus 4 (Non-reasoning): benchmark results
Claude Opus 4 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-05-22. Access: API.
Unified ELO 1590 ± 1, rank #478 of 1761 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Writing | 68.24 | Writing Score | 97.6 |
| Generalization V1 (Lechmazur) | 1.7 | Avg Rank (lower is better) | 95.6 |
| Confabulation Leaderboard (Lechmazur) | 2.97 | Confabulation rate % (lower is better) | 95.2 |
| UGI - Natural Intelligence | 57.57 | NatInt Score | 93 |
| Artificial Analysis Intelligence Index | 19.06 | Intelligence Index | 64.5 |
| Elimination Game (Lechmazur) | 4.41 | TrueSkill μ | 61 |
| NYT Connections Older Models | 19.7 | Score (%) | 60.5 |
| LLM Emergent Collusion | 36 | Collusion Rate (%) | 58.3 |
| Chess Bench LLM | 309 | Lichess Rating | 51.9 |
| AA GPQA Diamond | 70.1 | Accuracy (%) | 49 |
| AA IFBench | 43.27 | Accuracy (%) | 47.1 |
| UGI Leaderboard | 33.8 | UGI Score | 47 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.