Grok 4.20 (Reasoning): benchmark results

Grok 4.20 evaluated with reasoning enabled. Provider: xAI. Released 2026-03-10. Access: API.

Unified ELO 1655 ± 1, rank #227 of 1761 rated models, from 52 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Causal Sensitivity Score (CSS)47.3CSS (self-reported)100
SpeechMap Compliance98.2% Requests Completed99.1
Wolfram LLM Benchmarking Project66.3Correct Functionality (%)94.6
BenchTable76.8Total Score (%)94.2
AI Chess Leaderboard (Reasoning)1522Elo93.2
AI for Education SEND83.94Accuracy (%)91.6
Kagi LLM Benchmark75Accuracy (%)91.5
CLBench22.2Solving Rate (%)88.6
AI for Education Pedagogy - Science91.8Accuracy (%)88.4
AI for Education Pedagogy - Maths88.89Accuracy (%)85.5
AI Chess Leaderboard (Continuation)1087Elo84.8
LLM Chess (Saplin)843ELO84.5

Interactive version: theaggregate.ai/model?slug=grok-4-20-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.