Grok 4.3 — benchmark results

xAI's flagship Grok 4.3 multimodal model with native video input and a 1M-token context. Provider: xAI. Released 2026-04-30. Access: API.

Unified ELO 1701 ± 15, rank #267 of 1776 rated models, from 148 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BLXBench85.5Score (self-reported)100
Vals AI CaseLaw v279.31Accuracy (%)100
AI for Education Pedagogy - Technology88.68Accuracy (%)98.9
UGI Leaderboard59.41UGI Score98.7
Vals AI CorpFin v268.53Accuracy (%)96.7
Sycophancy (Lechmazur)0.5Sycophancy rate % (lower is better)95.2
AI Chess Leaderboard (Reasoning)1432Elo93
AI for Education Pedagogy - Secondary88.21Accuracy (%)92.7
UGI - Writing57.74Writing Score92.6
CritPt8Accuracy (self-reported)92.5
UGI - Natural Intelligence53.75NatInt Score92.1
AI for Education Pedagogy88.65Accuracy (%)92

Interactive version: theaggregate.ai/model?slug=grok-4-3 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.