Grok 4 — benchmark results
xAI's flagship Grok 4 reasoning model with native tool use. Provider: xAI. Released 2025-07-09. Access: API.
Unified ELO 1729 ± 10, rank #221 of 1776 rated models, from 399 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Balrog | 43.6 | Score (self-reported) | 100 |
| FinSearchComp | 68.9 | Average Score (%) | 100 |
| LLM Emergent Collusion | 75 | Collusion Rate (%) | 100 |
| LMGame-Bench Tetris | 125.7 | Score | 100 |
| Medmarks - M-ARC | 80 | Score (%) | 100 |
| Medmarks - MedConceptsQA Hard | 99.53 | Score (%) | 100 |
| Medmarks - MedConceptsQA Medium | 99.82 | Score (%) | 100 |
| OpenAI GPT-5 System Card - Molecular Biology Capabilities Test | 51.7 | Mean Accuracy (%) | 100 |
| ProLLM - OpenBook Q&A | 94.4 | Score (%) | 100 |
| RubberDuckBench | 69.29 | Performance (%) | 100 |
| StereoTales | 304 | Emissions (self-reported) | 100 |
| Vending-Bench | 4694.15 | Net Worth ($) | 100 |
Interactive version: theaggregate.ai/model?slug=grok-4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.