PlanBench-V - Spatial Relationship Reasoning: leaderboard

Metric: Mean score (out of 2) on topological, ordinal and metric spatial relationship reasoning, reasoning level; answers to expert-curated questions on Chinese territorial spatial planning maps are scored 0 to 2 by a GPT-4o-mini judge (temperature 0) against structured reference answers with annotated critical points; objective items use exact match or semantic similarity; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1Qwen 3.6 Plus1.67
2Gemini 2.5 Pro1.44
3Kimi K2.61.41
4GPT-5.41.38
5Qwen 3.6 Flash1.35
6GPT-4o1.31
7Claude Opus 4.71.3
8Qwen 2.5 VL 7B Instruct0.87
9InternVL3-8B0.8
10InternVL3-14B0.79
11GPT-4o Mini0.79
12Qwen 2 VL 7B0.72
13Qwen 2 VL 2B0.5

Interactive version: theaggregate.ai/benchmark?slug=planbench-v-spatial-relationship-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.