Grok 4.20 0309 (Reasoning): benchmark results
March 9, 2026 Grok 4.20 snapshot evaluated with reasoning enabled. Provider: xAI. Released 2026-03-09. Access: API.
Unified ELO 1629 ± 1, rank #322 of 1761 rated models, from 126 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 82.93 | Accuracy (%) | 99.8 |
| UGI Leaderboard | 64.23 | UGI Score | 99.4 |
| AA TAU-2 Bench | 96.49 | Accuracy (%) | 97 |
| AA Global-MMLU-Lite - Burmese | 88.83 | Accuracy (%) | 95.5 |
| Vals AI AIME | 96.46 | Accuracy (%) | 94.7 |
| AA Global-MMLU-Lite - Italian | 92.75 | Accuracy (%) | 94.6 |
| AA Global-MMLU-Lite - Portuguese | 92.33 | Accuracy (%) | 93.2 |
| UGI - Natural Intelligence | 52.75 | NatInt Score | 90.9 |
| UGI - Writing | 55.26 | Writing Score | 90.9 |
| AA Global-MMLU-Lite - Hindi | 89.92 | Accuracy (%) | 90.1 |
| Nejumi 4 - GLP - Abstract Reasoning | 81 | Score (%) | 89.3 |
| AA Global-MMLU-Lite - English | 93.42 | Accuracy (%) | 89.2 |
Interactive version: theaggregate.ai/model?slug=grok-4-20-0309-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.