Grok 4.3 (Medium) — benchmark results
Provider: xAI. Released 2026-04-30. Access: API.
Unified ELO 1764 ± 22, rank #208 of 1806 rated models, from 34 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 83.33 | Accuracy (%) | 100 |
| AA Omniscience | 16.7 | Score | 94.6 |
| AA GPQA Diamond | 88.99 | Accuracy (%) | 90.2 |
| AA Humanity's Last Exam | 29.98 | Accuracy (%) | 86.6 |
| AA TAU-2 Bench | 91.23 | Accuracy (%) | 86.1 |
| AA Omniscience - Software Engineering (SWE) - C | 69 | Accuracy (%) | 84.8 |
| Chess Bench LLM | 1248 | Lichess Rating | 84.5 |
| AA CritPt | 4.86 | Accuracy (%) | 83.7 |
| Epoch AI - Scicode | 44.56 | Score | 83.3 |
| Artificial Analysis Intelligence Index | 36.92 | Intelligence Index | 83.2 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 51.11 | Accuracy (%) | 81.7 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 52.73 | Accuracy (%) | 81.6 |
Interactive version: theaggregate.ai/model?slug=grok-4-3-medium · How It Works · Data refreshed daily, snapshot 2026-08-07.