Grok 4.20 0309 v2 (Non-reasoning): benchmark results

Provider: xAI. Released 2026-03-09. Access: API.

Unified ELO 1571 ± 1, rank #552 of 1761 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CritPt30Accuracy (self-reported)80.8
AA Humanity's Last Exam27.94Accuracy (%)79.5
AA Omniscience - Law23Accuracy (%)74.5
AA Omniscience - Humanities & Social Sciences27.4Accuracy (%)69.8
AA Omniscience - Health26Accuracy (%)69.5
AA Omniscience - Business22.4Accuracy (%)69.3
AA-Omniscience Accuracy26.68Accuracy (%)67.7
AA Omniscience - Software Engineering (SWE)35.2Accuracy (%)67.1
AA GPQA Diamond77.58Accuracy (%)62.5
Artificial Analysis Intelligence Index15.58Intelligence Index59.4
AA IFBench49.32Accuracy (%)58.8
AA TAU-2 Bench59.94Accuracy (%)56.6

Interactive version: theaggregate.ai/model?slug=grok-4-20-0309-v2-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.