Spatial Competence Benchmark - Axiomatic: leaderboard
Metric: Accuracy (%) over the 75 axiomatic subtasks (topology enumeration, edge enumeration, behaviour classification, half-subdivision neighbours, two-segment partitions) of the Spatial Competence Benchmark (SCBench) subtasks with executable outputs checked by deterministic verifiers or simulators, each subtask scored in [0, 1] (binary or partial credit), every subtask weighted equally; single-turn, zero-shot, one sample, no tools, no self-correction; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 81.3 |
| 2 | GPT-5.2 (xHigh) | 74.7 |
| 3 | Claude Sonnet 4.5 (Thinking) | 49.3 |
Interactive version: theaggregate.ai/benchmark?slug=spatial-competence-benchmark-axiomatic · How It Works · Data refreshed daily, snapshot 2026-10-07.