UAV-DualCog - Landmark-Driven Action Decision: leaderboard

Metric: Answer accuracy (%; 1,024 questions asking which way the UAV should move to approach a target landmark; structured JSON answers on renders of simulated AerialVLN scenes, the discrete answer scored whatever the predicted box or interval; instant mode with explicit thinking disabled where a model allows it; higher is better). Source: arxiv.org. Saturation forecast: Around 2031. 36 models tracked.

Top models

#ModelScore
1GPT-5.3 Instant62.3
2Claude Sonnet 4.661
3Qwen 3.5 35B A3B (Non-reasoning)57.5
4GPT-5.5 (Non-reasoning)53.6
5Qwen 3.5 9B (Non-reasoning)52.2
6Qwen 3.5 27B (Non-reasoning)50.6
7Qwen 3.5 122B A10B (Non-reasoning)49
8Qwen 3.5 Plus (Non-reasoning)48.9
9Qwen 3.5 4B (Non-reasoning)47.5
10Qwen 3.5 397B A17B (Non-reasoning)46.2
11GLM-4.6V (Non-reasoning)44.2
12Qwen 3.7 Plus (Non-reasoning)43.4
13Qwen 3.6 Plus (Non-reasoning)42.5
14MiMo-V2-Omni38.2
15GPT-5.4 (Non-reasoning)36

Interactive version: theaggregate.ai/benchmark?slug=uav-dualcog-landmark-driven-action-decision · How It Works · Data refreshed daily, snapshot 2026-09-29.