Grok 4: benchmark results

xAI's flagship Grok 4 reasoning model with native tool use. Provider: xAI. Released 2025-07-09. Access: API.

Unified ELO 1634 ± 1, rank #167 of 1392 rated models, from 467 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Balrog43.6Score (self-reported)100
FinSearchComp68.9Average Score (%)100
LMGame-Bench Tetris125.7Score100
Medmarks - M-ARC80Score (%)100
Medmarks - MedConceptsQA Hard99.53Score (%)100
Medmarks - MedConceptsQA Medium99.82Score (%)100
OpenAI GPT-5 System Card - Molecular Biology Capabilities Test51.7Mean Accuracy (%)100
ProLLM - OpenBook Q&A94.4Score (%)100
RubberDuckBench69.29Performance (%)100
Vending-Bench4694.15Net Worth ($)100
WDCD96.3DCD Score100
UGI Leaderboard67.83UGI Score99.8

Interactive version: theaggregate.ai/model?slug=grok-4 · How It Works · Data refreshed daily, snapshot 2026-09-05.