UAVQA-Bench - Quantity Awareness: leaderboard

Metric: Accuracy (%; scene and regional counting and number comparison; direct answering over UAVQA-Bench: 1,500 human-annotated questions on images from 13 public UAV datasets, 16 tasks in six capability dimensions, multiple choice plus box grounding counted correct at IoU 0.5). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro76.29
2Gemini 3 Flash71.71
3Qwen 3 VL 30B A3B (Thinking)69.43
4Qwen 3 VL 32B Instruct67.71
5Qwen 3 VL 32B (Thinking)67.14
6Qwen 3.5 9B65.14
7Qwen 3 VL 8B (Thinking)64.86
8Qwen 3 VL 8B Instruct64.57
9Qwen 3 VL 30B A3B Instruct58.57
10InternVL3.5-8B53.71

Interactive version: theaggregate.ai/benchmark?slug=uavqa-bench-quantity-awareness · How It Works · Data refreshed daily, snapshot 2026-09-29.