Claude Opus 4.1 (Non-reasoning) — benchmark results

Claude Opus 4.1 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-08-05. Access: API.

Unified ELO 1725 ± 33, rank #229 of 1776 rated models, from 11 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Generalization V1 (Lechmazur)1.69Avg Rank (lower is better)98.8
UGI - Writing69.91Writing Score98.8
UGI - Natural Intelligence58.74NatInt Score94.1
Confabulation Leaderboard (Lechmazur)3.96Confabulation rate % (lower is better)91.3
Elimination Game (Lechmazur)5.21TrueSkill μ84.7
Translation (Lechmazur)8.56Mean Score71.4
NYT Connections Older Models37.1Score (%)65.7
Chess Bench LLM367Lichess Rating48.5
Step Game (Lechmazur)1.57TrueSkill μ40.5
UGI Leaderboard31.94UGI Score38.6
UGI - Willingness (W/10)0.5W/10 Score0.9

Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.