Grok 4: benchmark results
xAI's flagship Grok 4 reasoning model with native tool use. Provider: xAI. Released 2025-07-09. Access: API.
Unified ELO 1634 ± 1, rank #167 of 1392 rated models, from 467 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Balrog | 43.6 | Score (self-reported) | 100 |
| FinSearchComp | 68.9 | Average Score (%) | 100 |
| LMGame-Bench Tetris | 125.7 | Score | 100 |
| Medmarks - M-ARC | 80 | Score (%) | 100 |
| Medmarks - MedConceptsQA Hard | 99.53 | Score (%) | 100 |
| Medmarks - MedConceptsQA Medium | 99.82 | Score (%) | 100 |
| OpenAI GPT-5 System Card - Molecular Biology Capabilities Test | 51.7 | Mean Accuracy (%) | 100 |
| ProLLM - OpenBook Q&A | 94.4 | Score (%) | 100 |
| RubberDuckBench | 69.29 | Performance (%) | 100 |
| Vending-Bench | 4694.15 | Net Worth ($) | 100 |
| WDCD | 96.3 | DCD Score | 100 |
| UGI Leaderboard | 67.83 | UGI Score | 99.8 |
Interactive version: theaggregate.ai/model?slug=grok-4 · How It Works · Data refreshed daily, snapshot 2026-09-05.