Grok 4.3 — benchmark results
xAI's flagship Grok 4.3 multimodal model with native video input and a 1M-token context. Provider: xAI. Released 2026-04-30. Access: API.
Unified ELO 1701 ± 15, rank #267 of 1776 rated models, from 148 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BLXBench | 85.5 | Score (self-reported) | 100 |
| Vals AI CaseLaw v2 | 79.31 | Accuracy (%) | 100 |
| AI for Education Pedagogy - Technology | 88.68 | Accuracy (%) | 98.9 |
| UGI Leaderboard | 59.41 | UGI Score | 98.7 |
| Vals AI CorpFin v2 | 68.53 | Accuracy (%) | 96.7 |
| Sycophancy (Lechmazur) | 0.5 | Sycophancy rate % (lower is better) | 95.2 |
| AI Chess Leaderboard (Reasoning) | 1432 | Elo | 93 |
| AI for Education Pedagogy - Secondary | 88.21 | Accuracy (%) | 92.7 |
| UGI - Writing | 57.74 | Writing Score | 92.6 |
| CritPt | 8 | Accuracy (self-reported) | 92.5 |
| UGI - Natural Intelligence | 53.75 | NatInt Score | 92.1 |
| AI for Education Pedagogy | 88.65 | Accuracy (%) | 92 |
Interactive version: theaggregate.ai/model?slug=grok-4-3 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.