Spatial Competence Benchmark: leaderboard

Metric: Accuracy (%) over all 285 subtasks of the Spatial Competence Benchmark (SCBench) subtasks with executable outputs checked by deterministic verifiers or simulators, each subtask scored in [0, 1] (binary or partial credit), every subtask weighted equally; single-turn, zero-shot, one sample, no tools, no self-correction; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 3 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)57.6
2GPT-5.2 (xHigh)57.6
3Claude Sonnet 4.5 (Thinking)34.9

Interactive version: theaggregate.ai/benchmark?slug=spatial-competence-benchmark · How It Works · Data refreshed daily, snapshot 2026-10-07.