Grok 4.6 (Medium): benchmark results
Provider: xAI. Released 2026-08-12. Access: API.
Unified ELO 1722 ± 1, rank #62 of 1761 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA GPQA Diamond | 93.54 | Accuracy (%) | 97.3 |
| Artificial Analysis Intelligence Index | 48.39 | Intelligence Index | 96.5 |
| AA Omniscience | 28 | Score | 95.5 |
| AA GDPval | 1647.49 | ELO | 94.9 |
| AA Humanity's Last Exam | 42.12 | Accuracy (%) | 92.8 |
| AA Long Context Reasoning | 81 | Accuracy (%) | 92.6 |
| Tau3 Banking | 44.33 | Success Rate (%) | 91.8 |
| AA CritPt | 17.71 | Accuracy (%) | 90.9 |
| AA Omniscience - Science, Engineering & Mathematics | 46.2 | Accuracy (%) | 89 |
| AA Omniscience - Law | 38.2 | Accuracy (%) | 87.5 |
| AA Omniscience - Software Engineering (SWE) | 64.9 | Accuracy (%) | 87 |
| AA-Omniscience Accuracy | 41.93 | Accuracy (%) | 85.4 |
Interactive version: theaggregate.ai/model?slug=grok-4-6-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.