Grok 2: benchmark results

xAI's second-generation Grok flagship (August 2024); weights opened in 2025, revealing a ~270B-total/~115B-active MoE. Provider: xAI. Released 2024-08-13. Access: API.

Unified ELO 1552 ± 1, rank #427 of 1392 rated models, from 39 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ProLLM - SQL Disambiguation47.9Score (%)96.1
Deception Effectiveness (Lechmazur)0.96Deception Score82.4
BELLS87.5BELLS Score (%)80
BenchTable61Total Score (%)72.2
LLM Stats (DocVQA)93.6Score (%)70.4
ProLLM - Function Calling89.1Score (%)68.5
AI Chess Leaderboard (Continuation)744Elo67
AI for Education Pedagogy - Social studies81.82Accuracy (%)63.9
AI for Education SEND77.06Accuracy (%)59.5
LLM Stats (MathVista)69Score (%)59.5
AI for Education Pedagogy - Technology80.19Accuracy (%)57.4
AI for Education Pedagogy - Secondary80.35Accuracy (%)56.7

Interactive version: theaggregate.ai/model?slug=grok-2 · How It Works · Data refreshed daily, snapshot 2026-09-05.