UAVQA-Bench - Spatial Relationship Understanding: leaderboard

Metric: Accuracy (%; spatial relations, height comparison and distance comparison from the aerial view; direct answering over UAVQA-Bench: 1,500 human-annotated questions on images from 13 public UAV datasets, 16 tasks in six capability dimensions, multiple choice plus box grounding counted correct at IoU 0.5). Source: arxiv.org. Saturation forecast: Around June 2027. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro60.8
2Qwen 3 VL 32B Instruct59.6
3Qwen 3 VL 32B (Thinking)58.4
4Gemini 3 Flash56.4
5Qwen 3 VL 30B A3B (Thinking)56
6Qwen 3.5 9B50.8
7InternVL3.5-8B47.2
8Qwen 3 VL 30B A3B Instruct43.2
9Qwen 3 VL 8B Instruct41.6
10Qwen 3 VL 8B (Thinking)39.6

Interactive version: theaggregate.ai/benchmark?slug=uavqa-bench-spatial-relationship-understanding · How It Works · Data refreshed daily, snapshot 2026-09-29.