MineBench: leaderboard
Evaluates LLM spatial reasoning through Minecraft-style voxel building tasks. Head-to-head comparisons judged by community votes. Tests 3D construction understanding from text prompts.
Metric: Elo Rating. Source: minebench.ai. Status: saturation imminent. 65 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 | 2111 |
| 2 | GPT-5.6 Pro Sol | 2108 |
| 3 | Claude Fable 5.1 | 2039 |
| 4 | GPT-5.5 Pro | 2028 |
| 5 | Claude Fable 5 | 1972 |
| 6 | GPT-5.5 | 1946 |
| 7 | Grok 4.6 | 1930 |
| 8 | Gemini 3.8 Flash | 1926 |
| 9 | GPT-5.4 Pro (xHigh) | 1892 |
| 10 | Claude Opus 4.8 | 1885 |
| 11 | GLM-5.3 Flash | 1878 |
| 12 | Gemini 3.7 Flash | 1869 |
| 13 | GLM-5.3 | 1852 |
| 14 | Muse Spark 1.3 | 1845 |
| 15 | Kimi K3 | 1811 |
Interactive version: theaggregate.ai/benchmark?slug=minebench · How It Works · Data refreshed daily, snapshot 2026-09-05.