PlanBench-V: leaderboard

Metric: Overall mean score (out of 2) on a 300-item stratified subset of the eight tasks; answers to expert-curated questions on Chinese territorial spatial planning maps are scored 0 to 2 by a GPT-4o-mini judge (temperature 0) against structured reference answers with annotated critical points; objective items use exact match or semantic similarity; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1Qwen 3.6 Plus1.7
2Gemini 2.5 Pro1.47
3GPT-5.41.43
4Kimi K2.61.42
5Claude Opus 4.71.38
6Qwen 3.6 Flash1.35
7GPT-4o1.34
8Qwen 2.5 VL 7B Instruct1.05
9InternVL3-14B0.98
10Qwen 2 VL 7B0.91
11InternVL3-8B0.91
12GPT-4o Mini0.87
13Qwen 2 VL 2B0.73

Interactive version: theaggregate.ai/benchmark?slug=planbench-v · How It Works · Data refreshed daily, snapshot 2026-09-29.