Spatial Competence Benchmark - Constructive: leaderboard
Metric: Accuracy (%) over the 171 constructive subtasks (CSG, projections, packing, triangulation, curve and loop fitting and similar construction tasks) of the Spatial Competence Benchmark (SCBench) subtasks with executable outputs checked by deterministic verifiers or simulators, each subtask scored in [0, 1] (binary or partial credit), every subtask weighted equally; single-turn, zero-shot, one sample, no tools, no self-correction; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 (xHigh) | 51.9 |
| 2 | Gemini 3 Pro (Preview) | 51.4 |
| 3 | Claude Sonnet 4.5 (Thinking) | 30.2 |
Interactive version: theaggregate.ai/benchmark?slug=spatial-competence-benchmark-constructive · How It Works · Data refreshed daily, snapshot 2026-10-07.