Grok 2 — benchmark results

xAI's second-generation Grok flagship (August 2024); weights opened in 2025, revealing a ~270B-total/~115B-active MoE. Provider: xAI. Released 2024-08-13. Access: API.

Unified ELO 1560 ± 18, rank #593 of 1776 rated models, from 33 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ProLLM - SQL Disambiguation47.9Score (%)96.1
MMVU63.4Score (self-reported)82.9
Deception Effectiveness (Lechmazur)0.96Deception Score82.4
BELLS87.5BELLS Score (%)80
BenchTable61Total Score (%)72.2
AI Chess Leaderboard (Continuation)744Elo69.1
ProLLM - Function Calling89.1Score (%)68.5
LLM Stats (DocVQA)93.6Score (%)68
AI for Education Pedagogy - Social studies81.82Accuracy (%)67.5
AI for Education SEND77.06Accuracy (%)63.2
AI for Education Pedagogy - Secondary80.35Accuracy (%)60.9
AI for Education Pedagogy - Technology80.19Accuracy (%)60.7

Interactive version: theaggregate.ai/model?slug=grok-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.