UAVQA-Bench - Visual Grounding: leaderboard

Metric: Accuracy (%; simple, complex-semantic and highest-object grounding with a normalized bounding box, correct at IoU 0.5 or more; direct answering over UAVQA-Bench: 1,500 human-annotated questions on images from 13 public UAV datasets, 16 tasks in six capability dimensions, multiple choice plus box grounding counted correct at IoU 0.5). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Flash82.55
2Gemini 3 Pro81.45
3Qwen 3 VL 8B Instruct80
4Qwen 3 VL 32B Instruct79.27
5Qwen 3 VL 32B (Thinking)79.27
6Qwen 3.5 9B78.55
7Qwen 3 VL 30B A3B Instruct75.64
8Qwen 3 VL 30B A3B (Thinking)50.91
9Qwen 3 VL 8B (Thinking)26.91
10InternVL3.5-8B3.63

Interactive version: theaggregate.ai/benchmark?slug=uavqa-bench-visual-grounding · How It Works · Data refreshed daily, snapshot 2026-09-29.