Grok 4 (High) — benchmark results

Provider: xAI. Released 2025-07-09. Access: API.

Unified ELO 1732 ± 56, rank #256 of 1839 rated models, from 8 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (HMMT25)96.7Score (%)100
Aider polyglot coding leaderboard79.6Pass rate (%)90.9
LLM Stats Score41.34LLM Stats Score (conservative rating)89.6
ZeroEval GPQA Diamond88.4GPQA Diamond Score87.9
ALL Bench Multimodal76.5Average Numeric VLM Score (%)46.7
NarrativeWorldBench78Plot-Beat F1 (h=50) (self-reported)38.9
ALL Bench LLM64.86Average Numeric Benchmark Score (%)34.2
FrontierMath - Tier 42.08Accuracy (%, 48 problems)23.9

Interactive version: theaggregate.ai/model?slug=grok-4-high · How It Works · Data refreshed daily, snapshot 2026-08-05.