Grok 4.5 (High) — benchmark results

Grok 4.5 evaluated at the high reasoning-effort setting. Provider: xAI. Released 2026-07-08. Access: API.

Unified ELO 1931 ± 21, rank #47 of 1776 rated models, from 51 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Omniscience26.38Score99.3
AA GPQA Diamond93.13Accuracy (%)99
Artificial Analysis Intelligence Index53.83Intelligence Index98.5
AA Omniscience - Software Engineering (SWE) - R70Accuracy (%)98.4
AA Omniscience - Health48.5Accuracy (%)98.2
AA Omniscience - Science, Engineering & Mathematics51.8Accuracy (%)98.2
Tau3 Banking32.58Success Rate (%)98.2
AA Humanity's Last Exam40.27Accuracy (%)97.4
AA SciCode54.05Accuracy (%)97.2
AA-Omniscience Accuracy52.05Accuracy (%)97.1
AA Omniscience - Humanities & Social Sciences50.2Accuracy (%)96.5
AA Omniscience - Law50.1Accuracy (%)96.4

Interactive version: theaggregate.ai/model?slug=grok-4-5-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.