Grok 4.20 (Reasoning) — benchmark results

Grok 4.20 evaluated with reasoning enabled. Provider: xAI. Released 2026-02-17. Access: API.

Unified ELO 1790 ± 32, rank #150 of 1776 rated models, from 46 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SpeechMap Compliance98.2% Requests Completed98.9
Wolfram LLM Benchmarking Project66.3Correct Functionality (%)95.8
AI Chess Leaderboard (Reasoning)1522Elo94.3
BenchTable76.8Total Score (%)94.2
AI for Education SEND83.94Accuracy (%)94.1
FutureX47.22Overall Score (latest week, %)92.9
LLM Chess (Saplin)843ELO91.4
AI for Education Pedagogy - Science91.8Accuracy (%)90.7
AI for Education Pedagogy - Maths88.89Accuracy (%)88.6
CLBench22.2Solving Rate (%)88.6
AI Chess Leaderboard (Continuation)1091Elo85.6
AI for Education Pedagogy - Primary91.08Accuracy (%)84.3

Interactive version: theaggregate.ai/model?slug=grok-4-20-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.