AirGroundBench - Localization Reasoning: leaderboard

Metric: Accuracy (%) on the localization spatial-reasoning VQA task; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 13 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro66.23
2Qwen 3.5 397B A17B57.12
3Qwen 3.5 Plus56.76
4GLM-4.6V56.76
5Claude Sonnet 4.656.36
6Qwen 3.5 27B56.32
7Kimi K2.555.41
8ERNIE 5.044.74
9GPT-5.243.24
10Gemini 3.1 Pro (Preview)41.11
11Qwen 3 VL 235B A22B40.54

Interactive version: theaggregate.ai/benchmark?slug=airgroundbench-localization-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.