Grok 4.3: benchmark results

xAI's flagship Grok 4.3 multimodal model with native video input and a 1M-token context. Provider: xAI. Released 2026-04-30. Access: API.

Unified ELO 1655 ± 1, rank #122 of 1392 rated models, from 243 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BLXBench85.5Score (self-reported)100
Vals AI CaseLaw v279.31Accuracy (%)100
UGI Leaderboard59.41UGI Score98.7
AI for Education Pedagogy - Technology88.68Accuracy (%)98.5
RAI-Bench - RAG Robustness (HY Abstention)91Rate (%)97.1
SpacetimeDB LLM Benchmark (Rust)97.8Eval Pass Rate (%)96.2
RAI-Bench - RAG Robustness (HY Factuality)48Rate (%)94.9
Vals AI CorpFin v268.53Accuracy (%)94.7
RAI-Bench - RAG Robustness (LC Abstention)93Rate (%)92
UGI - Writing57.74Writing Score91.9
AI Chess Leaderboard (Reasoning)1432Elo91.6
UGI - Natural Intelligence53.75NatInt Score91.4

Interactive version: theaggregate.ai/model?slug=grok-4-3 · How It Works · Data refreshed daily, snapshot 2026-09-05.