MineBench: leaderboard

Evaluates LLM spatial reasoning through Minecraft-style voxel building tasks. Head-to-head comparisons judged by community votes. Tests 3D construction understanding from text prompts.

Metric: Elo Rating. Source: minebench.ai. Status: saturation imminent. 65 models tracked.

Top models

#ModelScore
1Claude Opus 52111
2GPT-5.6 Pro Sol2108
3Claude Fable 5.12039
4GPT-5.5 Pro2028
5Claude Fable 51972
6GPT-5.51946
7Grok 4.61930
8Gemini 3.8 Flash1926
9GPT-5.4 Pro (xHigh)1892
10Claude Opus 4.81885
11GLM-5.3 Flash1878
12Gemini 3.7 Flash1869
13GLM-5.31852
14Muse Spark 1.31845
15Kimi K31811

Interactive version: theaggregate.ai/benchmark?slug=minebench · How It Works · Data refreshed daily, snapshot 2026-09-05.