Grok 4.20 0309 v2 (Non-reasoning): benchmark results
Provider: xAI. Released 2026-03-09. Access: API.
Unified ELO 1571 ± 1, rank #552 of 1761 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CritPt | 30 | Accuracy (self-reported) | 80.8 |
| AA Humanity's Last Exam | 27.94 | Accuracy (%) | 79.5 |
| AA Omniscience - Law | 23 | Accuracy (%) | 74.5 |
| AA Omniscience - Humanities & Social Sciences | 27.4 | Accuracy (%) | 69.8 |
| AA Omniscience - Health | 26 | Accuracy (%) | 69.5 |
| AA Omniscience - Business | 22.4 | Accuracy (%) | 69.3 |
| AA-Omniscience Accuracy | 26.68 | Accuracy (%) | 67.7 |
| AA Omniscience - Software Engineering (SWE) | 35.2 | Accuracy (%) | 67.1 |
| AA GPQA Diamond | 77.58 | Accuracy (%) | 62.5 |
| Artificial Analysis Intelligence Index | 15.58 | Intelligence Index | 59.4 |
| AA IFBench | 49.32 | Accuracy (%) | 58.8 |
| AA TAU-2 Bench | 59.94 | Accuracy (%) | 56.6 |
Interactive version: theaggregate.ai/model?slug=grok-4-20-0309-v2-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.