Grok 4.3 (High): benchmark results

Grok 4.3 evaluated at the high reasoning-effort setting. Provider: xAI. Released 2026-04-30. Access: API.

Unified ELO 1641 ± 1, rank #278 of 1761 rated models, from 63 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA IFBench81.29Accuracy (%)99.1
AA TAU-2 Bench97.66Accuracy (%)97.9
AA Omniscience17.98Score90.8
AA GPQA Diamond90.1Accuracy (%)89.6
AA Humanity's Last Exam37.21Accuracy (%)88.5
AA Terminal-Bench Hard37.88Accuracy (%)84.9
OTIS Mock AIME 2024-2593.33Accuracy (%)83.9
AA CritPt8Accuracy (%)82.6
Epoch AI - Dtbench90.67Score82.5
AA Omniscience - Science, Engineering & Mathematics42.38Accuracy (%)81.8
AA Omniscience - Health35.24Accuracy (%)80.9
Artificial Analysis Intelligence Index29.29Intelligence Index80.1

Interactive version: theaggregate.ai/model?slug=grok-4-3-high · How It Works · Data refreshed daily, snapshot 2026-09-05.