Grok Build 0.1: benchmark results

xAI's Grok Build 0.1 agentic coding model behind the Grok Build CLI. Provider: xAI. Released 2026-05-29. Access: API.

Unified ELO 1687 ± 1, rank #73 of 1392 rated models, from 66 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SpacetimeDB LLM Benchmark (Rust)97.8Eval Pass Rate (%)96.2
AA Omniscience - Health49Accuracy (%)95.5
AA Omniscience - Science, Engineering & Mathematics52.55Accuracy (%)95.3
AA Omniscience - Business43.84Accuracy (%)93.4
PinchBench92.07Success Rate (%)93.1
AA-Omniscience Accuracy51.52Accuracy (%)92.6
AA Omniscience - Law49.85Accuracy (%)92.2
AA Omniscience - Humanities & Social Sciences49.45Accuracy (%)92
AI Chess Leaderboard (Reasoning)1378Elo90.7
AA Humanity's Last Exam38.28Accuracy (%)89.5
AA GPQA Diamond89.49Accuracy (%)87.6
AA Omniscience6.37Score86.7

Interactive version: theaggregate.ai/model?slug=grok-build-0-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.