Grok 4 (High) — benchmark results
Provider: xAI. Released 2025-07-09. Access: API.
Unified ELO 1732 ± 56, rank #256 of 1839 rated models, from 8 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (HMMT25) | 96.7 | Score (%) | 100 |
| Aider polyglot coding leaderboard | 79.6 | Pass rate (%) | 90.9 |
| LLM Stats Score | 41.34 | LLM Stats Score (conservative rating) | 89.6 |
| ZeroEval GPQA Diamond | 88.4 | GPQA Diamond Score | 87.9 |
| ALL Bench Multimodal | 76.5 | Average Numeric VLM Score (%) | 46.7 |
| NarrativeWorldBench | 78 | Plot-Beat F1 (h=50) (self-reported) | 38.9 |
| ALL Bench LLM | 64.86 | Average Numeric Benchmark Score (%) | 34.2 |
| FrontierMath - Tier 4 | 2.08 | Accuracy (%, 48 problems) | 23.9 |
Interactive version: theaggregate.ai/model?slug=grok-4-high · How It Works · Data refreshed daily, snapshot 2026-08-05.