Grok 2 — benchmark results
xAI's second-generation Grok flagship (August 2024); weights opened in 2025, revealing a ~270B-total/~115B-active MoE. Provider: xAI. Released 2024-08-13. Access: API.
Unified ELO 1560 ± 18, rank #593 of 1776 rated models, from 33 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ProLLM - SQL Disambiguation | 47.9 | Score (%) | 96.1 |
| MMVU | 63.4 | Score (self-reported) | 82.9 |
| Deception Effectiveness (Lechmazur) | 0.96 | Deception Score | 82.4 |
| BELLS | 87.5 | BELLS Score (%) | 80 |
| BenchTable | 61 | Total Score (%) | 72.2 |
| AI Chess Leaderboard (Continuation) | 744 | Elo | 69.1 |
| ProLLM - Function Calling | 89.1 | Score (%) | 68.5 |
| LLM Stats (DocVQA) | 93.6 | Score (%) | 68 |
| AI for Education Pedagogy - Social studies | 81.82 | Accuracy (%) | 67.5 |
| AI for Education SEND | 77.06 | Accuracy (%) | 63.2 |
| AI for Education Pedagogy - Secondary | 80.35 | Accuracy (%) | 60.9 |
| AI for Education Pedagogy - Technology | 80.19 | Accuracy (%) | 60.7 |
Interactive version: theaggregate.ai/model?slug=grok-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.