Grok 4.20 Multi-Agent — benchmark results

xAI Grok 4.20 evaluated in multi-agent mode. Provider: xAI. Released 2026-03-09. Access: API.

Unified ELO 1828 ± 30, rank #117 of 1776 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI Leaderboard70UGI Score100
UGI - Writing63.13Writing Score95
FutureEval14.99Unified Forecasting Score94.6
Bullshit Benchmark67.3BS Detection Rate (%)93.8
UGI - Natural Intelligence56.34NatInt Score93
Chatbot Arena (Text)1471Elo92.8
AI Chess Leaderboard (Reasoning)1310Elo90.3
Chatbot Arena (Vision)1251Arena Score79.1
PM-LLM-Benchmark33.9Score77.4
Chatbot Arena (Search)1205Arena Score69.4
UGI - Willingness (W/10)6.5W/10 Score62.1
DystopiaBench68.1Dystopian Compliance Score (0-100)46.5

Interactive version: theaggregate.ai/model?slug=grok-4-20-multi-agent · How the rankings work · Data refreshed daily, snapshot 2026-07-22.