Grok 4.5: benchmark results
xAI's flagship Grok 4.5 model for reasoning and general tasks. Provider: xAI. Released 2026-07-08. Access: API.
Unified ELO 1715 ± 1, rank #35 of 1392 rated models, from 210 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BoundaryBench (NIST High) | 74.9 | Success Rate under NIST high policy (%) | 100 |
| E-commerce Last Exam - E-commerce | 78.65 | Mean Reward (%) | 100 |
| LLM Stats (Tau3 Banking) | 33 | Score (%) | 100 |
| Long-Horizon Terminal-Bench | 50.5 | Mean Score (%) | 100 |
| Nejumi 4 - ALT - Controllability | 94.73 | Score (%) | 100 |
| UGI Leaderboard | 62.32 | UGI Score | 99.1 |
| Conceptual Reasoning Index - Decision Theory (DTBench) | 94.22 | Chance-Corrected Score (0-100) | 97.9 |
| Kagi LLM Benchmark | 83.5 | Accuracy (%) | 96.5 |
| E-commerce Last Exam | 56.31 | Mean Reward (%) | 96.4 |
| E-commerce Last Exam - Travel | 46.73 | Mean Reward (%) | 96.4 |
| AI for Education Pedagogy - Secondary | 89.31 | Accuracy (%) | 96.2 |
| ZeroEval GPQA Diamond | 93 | GPQA Diamond Score | 96.2 |
Interactive version: theaggregate.ai/model?slug=grok-4-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.