Spatial Competence Benchmark - Planning: leaderboard

Metric: Accuracy (%) over the 39 planning subtasks (3D maze, hyper-snake, fluid simulation, terrain levelling) of the Spatial Competence Benchmark (SCBench) subtasks with executable outputs checked by deterministic verifiers or simulators, each subtask scored in [0, 1] (binary or partial credit), every subtask weighted equally; single-turn, zero-shot, one sample, no tools, no self-correction; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 3 models tracked.

Top models

#ModelScore
1GPT-5.2 (xHigh)50
2Gemini 3 Pro (Preview)39
3Claude Sonnet 4.5 (Thinking)27.5

Interactive version: theaggregate.ai/benchmark?slug=spatial-competence-benchmark-planning · How It Works · Data refreshed daily, snapshot 2026-10-07.