Grok 4.20 (Non-reasoning) — benchmark results
Grok 4.20 evaluated with reasoning disabled. Provider: xAI. Released 2026-03-10. Access: API.
Unified ELO 1634 ± 43, rank #408 of 1776 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Design Arena (Website) | 1240 | Elo | 72.4 |
| Design Arena (Data Viz) | 1234 | Elo | 71 |
| BenchTable | 58.7 | Total Score (%) | 68.6 |
| Design Arena (UI Components) | 1227 | Elo | 66.2 |
| Design Arena (Game Dev) | 1226 | Elo | 64.3 |
| Design Arena (3D) | 1207 | Elo | 62.9 |
| AI Chess Leaderboard (Reasoning) | 664 | Elo | 51.2 |
| CLBench | 17.7 | Solving Rate (%) | 45.7 |
| LLM Chess (Saplin) | -125 | ELO | 35.7 |
| Opper TaskBench | 82 | Avg Task Score (%) | 33.5 |
| AI Chess Leaderboard (Continuation) | 471 | Elo | 29.6 |
| Design Arena (SVG) | 1103 | Elo | 26.3 |
Interactive version: theaggregate.ai/model?slug=grok-4-20-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.