Spatial Competence Benchmark - Axiomatic: leaderboard

Metric: Accuracy (%) over the 75 axiomatic subtasks (topology enumeration, edge enumeration, behaviour classification, half-subdivision neighbours, two-segment partitions) of the Spatial Competence Benchmark (SCBench) subtasks with executable outputs checked by deterministic verifiers or simulators, each subtask scored in [0, 1] (binary or partial credit), every subtask weighted equally; single-turn, zero-shot, one sample, no tools, no self-correction; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 3 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)81.3
2GPT-5.2 (xHigh)74.7
3Claude Sonnet 4.5 (Thinking)49.3

Interactive version: theaggregate.ai/benchmark?slug=spatial-competence-benchmark-axiomatic · How It Works · Data refreshed daily, snapshot 2026-10-07.