Claude Opus 4.1 (Non-reasoning) — benchmark results
Claude Opus 4.1 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-08-05. Access: API.
Unified ELO 1725 ± 33, rank #229 of 1776 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Generalization V1 (Lechmazur) | 1.69 | Avg Rank (lower is better) | 98.8 |
| UGI - Writing | 69.91 | Writing Score | 98.8 |
| UGI - Natural Intelligence | 58.74 | NatInt Score | 94.1 |
| Confabulation Leaderboard (Lechmazur) | 3.96 | Confabulation rate % (lower is better) | 91.3 |
| Elimination Game (Lechmazur) | 5.21 | TrueSkill μ | 84.7 |
| Translation (Lechmazur) | 8.56 | Mean Score | 71.4 |
| NYT Connections Older Models | 37.1 | Score (%) | 65.7 |
| Chess Bench LLM | 367 | Lichess Rating | 48.5 |
| Step Game (Lechmazur) | 1.57 | TrueSkill μ | 40.5 |
| UGI Leaderboard | 31.94 | UGI Score | 38.6 |
| UGI - Willingness (W/10) | 0.5 | W/10 Score | 0.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.