Grok 4.20 (Reasoning): benchmark results
Grok 4.20 evaluated with reasoning enabled. Provider: xAI. Released 2026-03-10. Access: API.
Unified ELO 1655 ± 1, rank #227 of 1761 rated models, from 52 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Causal Sensitivity Score (CSS) | 47.3 | CSS (self-reported) | 100 |
| SpeechMap Compliance | 98.2 | % Requests Completed | 99.1 |
| Wolfram LLM Benchmarking Project | 66.3 | Correct Functionality (%) | 94.6 |
| BenchTable | 76.8 | Total Score (%) | 94.2 |
| AI Chess Leaderboard (Reasoning) | 1522 | Elo | 93.2 |
| AI for Education SEND | 83.94 | Accuracy (%) | 91.6 |
| Kagi LLM Benchmark | 75 | Accuracy (%) | 91.5 |
| CLBench | 22.2 | Solving Rate (%) | 88.6 |
| AI for Education Pedagogy - Science | 91.8 | Accuracy (%) | 88.4 |
| AI for Education Pedagogy - Maths | 88.89 | Accuracy (%) | 85.5 |
| AI Chess Leaderboard (Continuation) | 1087 | Elo | 84.8 |
| LLM Chess (Saplin) | 843 | ELO | 84.5 |
Interactive version: theaggregate.ai/model?slug=grok-4-20-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.