Grok 4.20 Multi-Agent — benchmark results
xAI Grok 4.20 evaluated in multi-agent mode. Provider: xAI. Released 2026-03-09. Access: API.
Unified ELO 1828 ± 30, rank #117 of 1776 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI Leaderboard | 70 | UGI Score | 100 |
| UGI - Writing | 63.13 | Writing Score | 95 |
| FutureEval | 14.99 | Unified Forecasting Score | 94.6 |
| Bullshit Benchmark | 67.3 | BS Detection Rate (%) | 93.8 |
| UGI - Natural Intelligence | 56.34 | NatInt Score | 93 |
| Chatbot Arena (Text) | 1471 | Elo | 92.8 |
| AI Chess Leaderboard (Reasoning) | 1310 | Elo | 90.3 |
| Chatbot Arena (Vision) | 1251 | Arena Score | 79.1 |
| PM-LLM-Benchmark | 33.9 | Score | 77.4 |
| Chatbot Arena (Search) | 1205 | Arena Score | 69.4 |
| UGI - Willingness (W/10) | 6.5 | W/10 Score | 62.1 |
| DystopiaBench | 68.1 | Dystopian Compliance Score (0-100) | 46.5 |
Interactive version: theaggregate.ai/model?slug=grok-4-20-multi-agent · How the rankings work · Data refreshed daily, snapshot 2026-07-22.