Claude Opus 4 (Non-reasoning) — benchmark results

Claude Opus 4 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-05-22. Access: API.

Unified ELO 1675 ± 30, rank #311 of 1776 rated models, from 21 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI - Writing68.24Writing Score98.2
Generalization V1 (Lechmazur)1.7Avg Rank (lower is better)95.6
Confabulation Leaderboard (Lechmazur)2.97Confabulation rate % (lower is better)95.2
AA MMLU-Pro85.98Accuracy (%)93.8
UGI - Natural Intelligence57.57NatInt Score93.7
AA SciCode40.86Accuracy (%)76.9
AA MATH-50094.07Accuracy (%)76.5
Artificial Analysis Intelligence Index25.5Intelligence Index69.3
NYT Connections Older Models34.4Score (%)62
Elimination Game (Lechmazur)4.41TrueSkill μ61
AA LiveCodeBench54.18Pass@1 (%)60.5
AA GPQA Diamond70.1Accuracy (%)53.6

Interactive version: theaggregate.ai/model?slug=claude-opus-4-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.