Grok 4.5 (High): benchmark results

Grok 4.5 evaluated at the high reasoning-effort setting. Provider: xAI. Released 2026-07-08. Access: API.

Unified ELO 1695 ± 1, rank #119 of 1761 rated models, from 80 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Vals AI SkillsBench66.03Accuracy (%)100
SWE-rebench63.78Resolved (%)99
Epoch AI - Dtbench96.53Score97.5
AA GPQA Diamond93.13Accuracy (%)96.1
AA Omniscience - Health48.4Accuracy (%)95.1
AA Omniscience - Science, Engineering & Mathematics52.32Accuracy (%)95.1
AA Omniscience25.32Score94.3
AA Humanity's Last Exam42.68Accuracy (%)94.1
Artificial Analysis Intelligence Index45.48Intelligence Index94.1
OTIS Mock AIME 2024-2597.78Accuracy (%)93.6
Vals AI Harvey Legal Agent Bench12.92Accuracy (%)93
AA-Omniscience Accuracy51.55Accuracy (%)92.8

Interactive version: theaggregate.ai/model?slug=grok-4-5-high · How It Works · Data refreshed daily, snapshot 2026-09-05.