MineBench — leaderboard

Evaluates LLM spatial reasoning through Minecraft-style voxel building tasks. Head-to-head comparisons judged by community votes. Tests 3D construction understanding from text prompts.

Metric: Elo Rating. Source: minebench.ai. Status: saturation imminent. 51 models tracked.

Top models

#ModelScore
1GPT-5.5 Pro1991.84
2Claude Fable 51904.6
3GPT-5.51897.42
4Kimi K31822.73
5Claude Opus 4.81814.95
6GPT-5.4 Pro (xHigh)1773.18
7Claude Sonnet 51638.24
8Gemini 3.5 Flash1636.54
9Gemini 3.1 Pro (Preview)1630.07
10GPT-5.2 Pro1608.29
11GPT-5.41605.82
12Grok 4.51578.25
13GPT-5.3 Codex1567
14Claude Opus 4.71555.88
15Claude Opus 4.61499.52

Interactive version: theaggregate.ai/benchmark?slug=minebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.