Grok 4.20 0309 (Reasoning) — benchmark results
March 9, 2026 Grok 4.20 snapshot evaluated with reasoning enabled. Provider: xAI. Released 2026-03-09. Access: API.
Unified ELO 1792 ± 17, rank #145 of 1776 rated models, from 90 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 82.93 | Accuracy (%) | 99.8 |
| UGI Leaderboard | 64.23 | UGI Score | 99.4 |
| AA TAU-2 Bench | 96.49 | Accuracy (%) | 97 |
| AA Global-MMLU-Lite - Burmese | 88.83 | Accuracy (%) | 95.5 |
| AA Global-MMLU-Lite - Italian | 92.75 | Accuracy (%) | 94.8 |
| Vals AI AIME | 96.46 | Accuracy (%) | 94.7 |
| AA Omniscience | 13.37 | Score | 93.8 |
| AA Global-MMLU-Lite - Portuguese | 92.33 | Accuracy (%) | 93.3 |
| UGI - Natural Intelligence | 52.75 | NatInt Score | 91.6 |
| UGI - Writing | 55.26 | Writing Score | 91.6 |
| AA GPQA Diamond | 88.48 | Accuracy (%) | 90.6 |
| AA Global-MMLU-Lite - Hindi | 89.92 | Accuracy (%) | 90.5 |
Interactive version: theaggregate.ai/model?slug=grok-4-20-0309-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.