SnorkelSpatial: leaderboard

Metric: Accuracy@1 (%, 330 spatial reasoning questions). Source: snorkel.ai. 33 models tracked.

Top models

#ModelScore
1GPT-5.499
2Grok 4 Fast (Reasoning)84.85
3O376.67
4GPT-573.94
5GPT-OSS-120B52.73
6GPT-5 Mini45.45
7Claude Opus 4.145.15
8Magistral Medium 1.244.24
9Claude Opus 440.3
10O3 Mini37.88
11Claude Sonnet 433.33
12GPT-5 Nano26.67
13Claude 3.7 Sonnet21.52
14Gemini 2.5 Flash18.79
15Llama 4 Scout15.45

Interactive version: theaggregate.ai/benchmark?slug=snorkelspatial · How It Works · Data refreshed daily, snapshot 2026-09-19.