Grok 4.3 (High) — benchmark results
Grok 4.3 evaluated at the high reasoning-effort setting. Provider: xAI. Released 2026-04-30. Access: API.
Unified ELO 1824 ± 26, rank #119 of 1776 rated models, from 51 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 81.29 | Accuracy (%) | 99.1 |
| AA TAU-2 Bench | 97.66 | Accuracy (%) | 97.9 |
| AA Omniscience | 18.32 | Score | 96.4 |
| AA GPQA Diamond | 90.1 | Accuracy (%) | 94.5 |
| AA Humanity's Last Exam | 35.03 | Accuracy (%) | 93.1 |
| AA Omniscience - Software Engineering (SWE) - C | 74 | Accuracy (%) | 90.5 |
| AA SciCode | 47.34 | Accuracy (%) | 89.7 |
| AA CritPt | 8 | Accuracy (%) | 89.2 |
| AA Omniscience - Science, Engineering & Mathematics | 42.3 | Accuracy (%) | 88.7 |
| Artificial Analysis Intelligence Index | 37.58 | Intelligence Index | 86.9 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 58.18 | Accuracy (%) | 86.2 |
| AA Omniscience - Software Engineering (SWE) - Rust | 70 | Accuracy (%) | 85.7 |
Interactive version: theaggregate.ai/model?slug=grok-4-3-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.