UAV-DualCog - Self-Relative Position: leaderboard

Metric: Answer accuracy (%; 1,024 questions asking where the UAV is relative to a landmark from 2 or 5 aerial images; structured JSON answers on renders of simulated AerialVLN scenes, the discrete answer scored whatever the predicted box or interval; instant mode with explicit thinking disabled where a model allows it; higher is better). Source: arxiv.org. Saturation forecast: Around 2032. 36 models tracked.

Top models

#ModelScore
1GPT-5.5 (Non-reasoning)38
2GPT-5.4 (Non-reasoning)37.9
3Claude Sonnet 4.637.7
4Kimi K2.6 (Non-reasoning)36.1
5GPT-5.3 Instant35.2
6Kimi K2.5 (Non-reasoning)34.3
7Gemini 3.1 Flash Lite (Non-reasoning)33.7
8Qwen 3.7 Plus (Non-reasoning)33.3
9Qwen 3.6 Plus (Non-reasoning)32.8
10Qwen 3.5 122B A10B (Non-reasoning)32.6
11Qwen 3.5 27B (Non-reasoning)31.6
12Qwen 3.5 4B (Non-reasoning)30.9
13GLM-4.6V (Non-reasoning)30.8
14Qwen 3.5 35B A3B (Non-reasoning)29.5
15Qwen 3.5 9B (Non-reasoning)29

Interactive version: theaggregate.ai/benchmark?slug=uav-dualcog-self-relative-position · How It Works · Data refreshed daily, snapshot 2026-09-29.