Grok 4.20 (High): benchmark results
Provider: xAI. Released 2026-03-10. Access: API.
Unified ELO 1687 ± 18, rank #382 of 2088 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CLBench Life - Personal Information Fragments | 13.3 | Task solving rate (%) on the 45 personal-information-fragmen | 66.7 |
| MathDuels - Solver Rating | 1950 | Solver rating (Elo-style points, open scale): Rasch ability | 66.7 |
| CLBench Life - Fragmented Information & Revisions | 11.9 | Solving Rate (%) | 62.5 |
| CLBench Life - Digital Footprints & Daily-Life Records | 13.3 | Task solving rate (%) on the 45 digital-footprint and daily- | 61.1 |
| CLBench Life - Public Information Fragments | 12.6 | Task solving rate (%) on the 45 public-information-fragment | 55.6 |
| CLBench Life - Behavioral Records & Activity Trails | 10.9 | Solving Rate (%) | 46.4 |
| CLBench Life | 11.9 | Solving Rate (%) | 42.9 |
| CLBench Life - Communication & Social Interactions | 12.8 | Solving Rate (%) | 42.9 |
| K-Bench (Therapeutic v0) | 97.24 | Overall (%, Therapeutic v0 prompt) | 41.9 |
| CLBench Life - Community Interactions | 10.4 | Task solving rate (%) on the 45 community-interaction (forum | 38.9 |
| MathDuels | 1485 | Composite rating (Elo-style points, open scale): the mean of | 38.9 |
| MazeBench | 0 | Gems collected (gems) | 33.8 |
Interactive version: theaggregate.ai/model?slug=grok-4-20-high · How It Works · Data refreshed daily, snapshot 2026-10-07.