Grok 4.6: benchmark results
Provider: xAI. Released 2026-08-12. Access: API.
Unified ELO 1741 ± 1, rank #19 of 1392 rated models, from 113 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI for Education Pedagogy - Science | 96.72 | Accuracy (%) | 100 |
| AI for Education Pedagogy - Secondary | 91.19 | Accuracy (%) | 100 |
| Benchmarks.bio - BioSecBench-Refusal | 64.2 | Pass Rate (%) | 100 |
| LLM Stats (APEX-Agents) | 57.5 | Score (%) | 100 |
| MERA v2 - Agentic | 85.5 | Score (%) | 100 |
| MERA v2 - Enantiosemy | 70.8 | Score (%) | 100 |
| MERA v2 - GorillaHard | 76.2 | Score (%) | 100 |
| AI for Education Pedagogy | 91.55 | Accuracy (%) | 99.4 |
| Conceptual Reasoning Index - Decision Theory (DTBench) | 95.55 | Chance-Corrected Score (0-100) | 99 |
| OpenUI Generative UI Benchmark | 99.5 | Structural Validity (%) | 98.3 |
| RuneBench | 12757 | Total Peak XP Rate (XP/min) | 98.1 |
| BenchmarkList ECI | 151.28 | Capability Index (ECI) | 98 |
Interactive version: theaggregate.ai/model?slug=grok-4-6 · How It Works · Data refreshed daily, snapshot 2026-09-05.