Claude Opus 4.1 (Non-reasoning): benchmark results
Claude Opus 4.1 evaluated with reasoning disabled. Provider: Anthropic. Released 2025-08-05. Access: API.
Unified ELO 1624 ± 1, rank #341 of 1761 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Generalization V1 (Lechmazur) | 1.69 | Avg Rank (lower is better) | 98.8 |
| UGI - Writing | 69.91 | Writing Score | 98.3 |
| UGI - Natural Intelligence | 58.74 | NatInt Score | 93.4 |
| Confabulation Leaderboard (Lechmazur) | 3.96 | Confabulation rate % (lower is better) | 91.3 |
| Elimination Game (Lechmazur) | 5.21 | TrueSkill μ | 84.7 |
| Translation (Lechmazur) | 8.56 | Mean Score | 71.4 |
| Artificial Analysis Intelligence Index | 21.67 | Intelligence Index | 67.5 |
| NYT Connections Older Models | 21.7 | Score (%) | 64.5 |
| Chess Bench LLM | 315 | Lichess Rating | 52.8 |
| Step Game (Lechmazur) | 1.57 | TrueSkill μ | 40.5 |
| UGI Leaderboard | 31.94 | UGI Score | 38.2 |
| UGI - Willingness (W/10) | 0.5 | W/10 Score | 0.9 |
Interactive version: theaggregate.ai/model?slug=claude-opus-4-1-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.