Grok Build 0.1: benchmark results
xAI's Grok Build 0.1 agentic coding model behind the Grok Build CLI. Provider: xAI. Released 2026-05-29. Access: API.
Unified ELO 1687 ± 1, rank #73 of 1392 rated models, from 66 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SpacetimeDB LLM Benchmark (Rust) | 97.8 | Eval Pass Rate (%) | 96.2 |
| AA Omniscience - Health | 49 | Accuracy (%) | 95.5 |
| AA Omniscience - Science, Engineering & Mathematics | 52.55 | Accuracy (%) | 95.3 |
| AA Omniscience - Business | 43.84 | Accuracy (%) | 93.4 |
| PinchBench | 92.07 | Success Rate (%) | 93.1 |
| AA-Omniscience Accuracy | 51.52 | Accuracy (%) | 92.6 |
| AA Omniscience - Law | 49.85 | Accuracy (%) | 92.2 |
| AA Omniscience - Humanities & Social Sciences | 49.45 | Accuracy (%) | 92 |
| AI Chess Leaderboard (Reasoning) | 1378 | Elo | 90.7 |
| AA Humanity's Last Exam | 38.28 | Accuracy (%) | 89.5 |
| AA GPQA Diamond | 89.49 | Accuracy (%) | 87.6 |
| AA Omniscience | 6.37 | Score | 86.7 |
Interactive version: theaggregate.ai/model?slug=grok-build-0-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.