UAVQA-Bench - Fine-Grained Attribute Perception: leaderboard

Metric: Accuracy (%; attribute, function and safe-landing recognition; direct answering over UAVQA-Bench: 1,500 human-annotated questions on images from 13 public UAV datasets, 16 tasks in six capability dimensions, multiple choice plus box grounding counted correct at IoU 0.5). Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro68.5
2Gemini 3 Flash60.5
3Qwen 3 VL 32B Instruct58
4Qwen 3.5 9B57.5
5Qwen 3 VL 32B (Thinking)57
6Qwen 3 VL 30B A3B (Thinking)53.5
7Qwen 3 VL 30B A3B Instruct51.5
8Qwen 3 VL 8B Instruct46.5
9Qwen 3 VL 8B (Thinking)44.5
10InternVL3.5-8B36.5

Interactive version: theaggregate.ai/benchmark?slug=uavqa-bench-fine-grained-attribute-perception · How It Works · Data refreshed daily, snapshot 2026-09-29.