Grok 4.5 — benchmark results

xAI's flagship Grok 4.5 model for reasoning and general tasks. Provider: xAI. Released 2026-07-08. Access: API.

Unified ELO 1882 ± 17, rank #77 of 1776 rated models, from 89 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Benchmarks.bio - BioSecBench-Surveillance53.3Pass Rate (%)100
Long-Horizon Terminal-Bench50.5Mean Score (%)100
Vals AI SkillsBench66.03Accuracy (%)100
UGI Leaderboard62.32UGI Score99.1
AI for Education Pedagogy - Secondary89.31Accuracy (%)97.7
BenchLM76.7Overall Score97
AI for Education Pedagogy89.66Accuracy (%)96.8
AI for Education Pedagogy - Science93.44Accuracy (%)96.6
ZeroEval GPQA Diamond93GPQA Diamond Score96.5
Kagi LLM Benchmark83.5Accuracy (%)96.4
Vals AI GPQA92.93Accuracy (%)96.3
Vals AI Harvey Legal Agent Bench12.92Accuracy (%)95.8

Interactive version: theaggregate.ai/model?slug=grok-4-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.