Grok 4.3 (High): benchmark results
Grok 4.3 evaluated at the high reasoning-effort setting. Provider: xAI. Released 2026-04-30. Access: API.
Unified ELO 1641 ± 1, rank #278 of 1761 rated models, from 63 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 81.29 | Accuracy (%) | 99.1 |
| AA TAU-2 Bench | 97.66 | Accuracy (%) | 97.9 |
| AA Omniscience | 17.98 | Score | 90.8 |
| AA GPQA Diamond | 90.1 | Accuracy (%) | 89.6 |
| AA Humanity's Last Exam | 37.21 | Accuracy (%) | 88.5 |
| AA Terminal-Bench Hard | 37.88 | Accuracy (%) | 84.9 |
| OTIS Mock AIME 2024-25 | 93.33 | Accuracy (%) | 83.9 |
| AA CritPt | 8 | Accuracy (%) | 82.6 |
| Epoch AI - Dtbench | 90.67 | Score | 82.5 |
| AA Omniscience - Science, Engineering & Mathematics | 42.38 | Accuracy (%) | 81.8 |
| AA Omniscience - Health | 35.24 | Accuracy (%) | 80.9 |
| Artificial Analysis Intelligence Index | 29.29 | Intelligence Index | 80.1 |
Interactive version: theaggregate.ai/model?slug=grok-4-3-high · How It Works · Data refreshed daily, snapshot 2026-09-05.