Claude Opus 4.1 (Non-reasoning): benchmark results

Claude Opus 4.1 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-08-05. Access: API.

Unified ELO 1624 ± 1, rank #341 of 1761 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V1 (Lechmazur)1.69Avg Rank (lower is better)98.8
UGI - Writing69.91Writing Score98.3
UGI - Natural Intelligence58.74NatInt Score93.4
Confabulation Leaderboard (Lechmazur)3.96Confabulation rate % (lower is better)91.3
Elimination Game (Lechmazur)5.21TrueSkill μ84.7
Translation (Lechmazur)8.56Mean Score71.4
Artificial Analysis Intelligence Index21.67Intelligence Index67.5
NYT Connections Older Models21.7Score (%)64.5
Chess Bench LLM315Lichess Rating52.8
Step Game (Lechmazur)1.57TrueSkill μ40.5
UGI Leaderboard31.94UGI Score38.2
UGI - Willingness (W/10)0.5W/10 Score0.9

Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.