Grok 4.6: benchmark results

Provider: xAI. Released 2026-08-12. Access: API.

Unified ELO 1741 ± 1, rank #19 of 1392 rated models, from 113 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI for Education Pedagogy - Science96.72Accuracy (%)100
AI for Education Pedagogy - Secondary91.19Accuracy (%)100
Benchmarks.bio - BioSecBench-Refusal64.2Pass Rate (%)100
LLM Stats (APEX-Agents)57.5Score (%)100
MERA v2 - Agentic85.5Score (%)100
MERA v2 - Enantiosemy70.8Score (%)100
MERA v2 - GorillaHard76.2Score (%)100
AI for Education Pedagogy91.55Accuracy (%)99.4
Conceptual Reasoning Index - Decision Theory (DTBench)95.55Chance-Corrected Score (0-100)99
OpenUI Generative UI Benchmark99.5Structural Validity (%)98.3
RuneBench12757Total Peak XP Rate (XP/min)98.1
BenchmarkList ECI151.28Capability Index (ECI)98

Interactive version: theaggregate.ai/model?slug=grok-4-6 · How It Works · Data refreshed daily, snapshot 2026-09-05.