UAVQA-Bench: leaderboard

Metric: Overall accuracy (%; sample-weighted accuracy over all 1,500 questions; direct answering over UAVQA-Bench: 1,500 human-annotated questions on images from 13 public UAV datasets, 16 tasks in six capability dimensions, multiple choice plus box grounding counted correct at IoU 0.5). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro73
2Gemini 3 Flash70.73
3Qwen 3 VL 32B Instruct69.6
4Qwen 3 VL 32B (Thinking)68.27
5Qwen 3.5 9B63.8
6Qwen 3 VL 30B A3B (Thinking)62
7Qwen 3 VL 8B Instruct61.73
8Qwen 3 VL 30B A3B Instruct61.13
9Qwen 3 VL 8B (Thinking)52.6
10InternVL3.5-8B40.53

Interactive version: theaggregate.ai/benchmark?slug=uavqa-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.