Grok 4.20 (High): benchmark results

Provider: xAI. Released 2026-03-10. Access: API.

Unified ELO 1687 ± 18, rank #382 of 2088 rated models, from 26 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CLBench Life - Personal Information Fragments13.3Task solving rate (%) on the 45 personal-information-fragmen66.7
MathDuels - Solver Rating1950Solver rating (Elo-style points, open scale): Rasch ability 66.7
CLBench Life - Fragmented Information & Revisions11.9Solving Rate (%)62.5
CLBench Life - Digital Footprints & Daily-Life Records13.3Task solving rate (%) on the 45 digital-footprint and daily-61.1
CLBench Life - Public Information Fragments12.6Task solving rate (%) on the 45 public-information-fragment 55.6
CLBench Life - Behavioral Records & Activity Trails10.9Solving Rate (%)46.4
CLBench Life11.9Solving Rate (%)42.9
CLBench Life - Communication & Social Interactions12.8Solving Rate (%)42.9
K-Bench (Therapeutic v0)97.24Overall (%, Therapeutic v0 prompt)41.9
CLBench Life - Community Interactions10.4Task solving rate (%) on the 45 community-interaction (forum38.9
MathDuels1485Composite rating (Elo-style points, open scale): the mean of38.9
MazeBench0Gems collected (gems)33.8

Interactive version: theaggregate.ai/model?slug=grok-4-20-high · How It Works · Data refreshed daily, snapshot 2026-10-07.