Spatial Competence Benchmark - Planning: leaderboard
Metric: Accuracy (%) over the 39 planning subtasks (3D maze, hyper-snake, fluid simulation, terrain levelling) of the Spatial Competence Benchmark (SCBench) subtasks with executable outputs checked by deterministic verifiers or simulators, each subtask scored in [0, 1] (binary or partial credit), every subtask weighted equally; single-turn, zero-shot, one sample, no tools, no self-correction; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 (xHigh) | 50 |
| 2 | Gemini 3 Pro (Preview) | 39 |
| 3 | Claude Sonnet 4.5 (Thinking) | 27.5 |
Interactive version: theaggregate.ai/benchmark?slug=spatial-competence-benchmark-planning · How It Works · Data refreshed daily, snapshot 2026-10-07.