Grok 4.3: benchmark results
xAI's flagship Grok 4.3 multimodal model with native video input and a 1M-token context. Provider: xAI. Released 2026-04-30. Access: API.
Unified ELO 1655 ± 1, rank #122 of 1392 rated models, from 243 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BLXBench | 85.5 | Score (self-reported) | 100 |
| Vals AI CaseLaw v2 | 79.31 | Accuracy (%) | 100 |
| UGI Leaderboard | 59.41 | UGI Score | 98.7 |
| AI for Education Pedagogy - Technology | 88.68 | Accuracy (%) | 98.5 |
| RAI-Bench - RAG Robustness (HY Abstention) | 91 | Rate (%) | 97.1 |
| SpacetimeDB LLM Benchmark (Rust) | 97.8 | Eval Pass Rate (%) | 96.2 |
| RAI-Bench - RAG Robustness (HY Factuality) | 48 | Rate (%) | 94.9 |
| Vals AI CorpFin v2 | 68.53 | Accuracy (%) | 94.7 |
| RAI-Bench - RAG Robustness (LC Abstention) | 93 | Rate (%) | 92 |
| UGI - Writing | 57.74 | Writing Score | 91.9 |
| AI Chess Leaderboard (Reasoning) | 1432 | Elo | 91.6 |
| UGI - Natural Intelligence | 53.75 | NatInt Score | 91.4 |
Interactive version: theaggregate.ai/model?slug=grok-4-3 · How It Works · Data refreshed daily, snapshot 2026-09-05.