Claude Opus 4 (Non-reasoning): benchmark results

Claude Opus 4 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1590 ± 1, rank #478 of 1761 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI - Writing68.24Writing Score97.6
Generalization V1 (Lechmazur)1.7Avg Rank (lower is better)95.6
Confabulation Leaderboard (Lechmazur)2.97Confabulation rate % (lower is better)95.2
UGI - Natural Intelligence57.57NatInt Score93
Artificial Analysis Intelligence Index19.06Intelligence Index64.5
Elimination Game (Lechmazur)4.41TrueSkill μ61
NYT Connections Older Models19.7Score (%)60.5
LLM Emergent Collusion36Collusion Rate (%)58.3
Chess Bench LLM309Lichess Rating51.9
AA GPQA Diamond70.1Accuracy (%)49
AA IFBench43.27Accuracy (%)47.1
UGI Leaderboard33.8UGI Score47

Interactive version: theaggregate.ai/model?slug=claude-opus-4-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.