UAVQA-Bench - Category Recognition: leaderboard

Metric: Accuracy (%; regional classification of detected entities; direct answering over UAVQA-Bench: 1,500 human-annotated questions on images from 13 public UAV datasets, 16 tasks in six capability dimensions, multiple choice plus box grounding counted correct at IoU 0.5). Source: arxiv.org. Saturation forecast: Around February 2027. 14 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B Instruct67.2
2Qwen 3 VL 32B (Thinking)63.2
3Gemini 3 Pro60.8
4Qwen 3.5 9B59.2
5Gemini 3 Flash59.2
6Qwen 3 VL 30B A3B Instruct54.4
7Qwen 3 VL 8B Instruct52.8
8Qwen 3 VL 30B A3B (Thinking)52
9Qwen 3 VL 8B (Thinking)50.4
10InternVL3.5-8B36

Interactive version: theaggregate.ai/benchmark?slug=uavqa-bench-category-recognition · How It Works · Data refreshed daily, snapshot 2026-09-29.