Grok 2: benchmark results
xAI's second-generation Grok flagship (August 2024); weights opened in 2025, revealing a ~270B-total/~115B-active MoE. Provider: xAI. Released 2024-08-13. Access: API.
Unified ELO 1552 ± 1, rank #427 of 1392 rated models, from 39 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ProLLM - SQL Disambiguation | 47.9 | Score (%) | 96.1 |
| Deception Effectiveness (Lechmazur) | 0.96 | Deception Score | 82.4 |
| BELLS | 87.5 | BELLS Score (%) | 80 |
| BenchTable | 61 | Total Score (%) | 72.2 |
| LLM Stats (DocVQA) | 93.6 | Score (%) | 70.4 |
| ProLLM - Function Calling | 89.1 | Score (%) | 68.5 |
| AI Chess Leaderboard (Continuation) | 744 | Elo | 67 |
| AI for Education Pedagogy - Social studies | 81.82 | Accuracy (%) | 63.9 |
| AI for Education SEND | 77.06 | Accuracy (%) | 59.5 |
| LLM Stats (MathVista) | 69 | Score (%) | 59.5 |
| AI for Education Pedagogy - Technology | 80.19 | Accuracy (%) | 57.4 |
| AI for Education Pedagogy - Secondary | 80.35 | Accuracy (%) | 56.7 |
Interactive version: theaggregate.ai/model?slug=grok-2 · How It Works · Data refreshed daily, snapshot 2026-09-05.