Grok 4.20 (Non-reasoning): benchmark results
Grok 4.20 evaluated with reasoning disabled. Provider: xAI. Released 2026-03-10. Access: API.
Unified ELO 1587 ± 1, rank #486 of 1761 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenchTable | 58.7 | Total Score (%) | 68.6 |
| Design Arena (Website) | 1235 | Elo | 66.8 |
| Design Arena (Data Viz) | 1220 | Elo | 65.1 |
| ZeroBench | 10 | Score (%) | 64.5 |
| Design Arena (UI Components) | 1209 | Elo | 59.7 |
| Design Arena (Game Dev) | 1205 | Elo | 58.8 |
| Design Arena (3D) | 1190 | Elo | 56.8 |
| AI Chess Leaderboard (Reasoning) | 664 | Elo | 47.7 |
| CLBench | 17.7 | Solving Rate (%) | 45.7 |
| Design Arena (ASCII Art) | 1152 | Elo | 36.6 |
| Opper TaskBench | 82 | Avg Task Score (%) | 33.5 |
| LLM Chess (Saplin) | -125 | ELO | 32.3 |
Interactive version: theaggregate.ai/model?slug=grok-4-20-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.