MineBench — leaderboard
Evaluates LLM spatial reasoning through Minecraft-style voxel building tasks. Head-to-head comparisons judged by community votes. Tests 3D construction understanding from text prompts.
Metric: Elo Rating. Source: minebench.ai. Status: saturation imminent. 51 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 Pro | 1991.84 |
| 2 | Claude Fable 5 | 1904.6 |
| 3 | GPT-5.5 | 1897.42 |
| 4 | Kimi K3 | 1822.73 |
| 5 | Claude Opus 4.8 | 1814.95 |
| 6 | GPT-5.4 Pro (xHigh) | 1773.18 |
| 7 | Claude Sonnet 5 | 1638.24 |
| 8 | Gemini 3.5 Flash | 1636.54 |
| 9 | Gemini 3.1 Pro (Preview) | 1630.07 |
| 10 | GPT-5.2 Pro | 1608.29 |
| 11 | GPT-5.4 | 1605.82 |
| 12 | Grok 4.5 | 1578.25 |
| 13 | GPT-5.3 Codex | 1567 |
| 14 | Claude Opus 4.7 | 1555.88 |
| 15 | Claude Opus 4.6 | 1499.52 |
Interactive version: theaggregate.ai/benchmark?slug=minebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.