Grok Build 0.1 — benchmark results

xAI's Grok Build 0.1 agentic coding model behind the Grok Build CLI. Provider: xAI. Released 2026-05-29. Access: API.

Unified ELO 1853 ± 25, rank #103 of 1776 rated models, from 50 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Omniscience - Health48.9Accuracy (%)99.1
AA Omniscience - Science, Engineering & Mathematics52.1Accuracy (%)99
AA Omniscience - Business44.8Accuracy (%)96.9
AA Omniscience - Humanities & Social Sciences50.2Accuracy (%)96.5
AA-Omniscience Accuracy51.4Accuracy (%)96.4
AA Omniscience - Law48.6Accuracy (%)95.8
AA Omniscience - Software Engineering (SWE) - C83Accuracy (%)95.1
AA Omniscience - Software Engineering (SWE) - Java55Accuracy (%)94.9
AA Humanity's Last Exam35.96Accuracy (%)94.1
AA Omniscience - Software Engineering (SWE) - R58Accuracy (%)93.4
AA Omniscience - Software Engineering (SWE) - JavaScript72.73Accuracy (%)93.2
AA SciCode50.23Accuracy (%)93.2

Interactive version: theaggregate.ai/model?slug=grok-build-0-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.