SIS-Bench: leaderboard

Metric: Overall accuracy (%; 4,856 questions over 13 tasks on UAV spatial cognition and self-awareness at perception, memory and reasoning levels, question-weighted; four-option multiple choice on UAV videos (at most 32 frames), zero-shot, open models through vLLM with at most 128 generated tokens and proprietary models through their APIs; higher is better). Source: arxiv.org. Saturation forecast: Around September 2027. 26 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)71.6
2Kimi K2.571
3Seed 1.870.6
4Qwen 3.5 Plus70.1
5GPT-5.470
6Qwen 3 VL 8B (Thinking)64.7
7Qwen 3 VL 8B Instruct64.6
8Qwen 3 VL 4B Instruct62.5
9InternVL3.5-8B61.1
10Qwen 3 VL 30B A3B Instruct61
11GLM-4.1V-9B (Thinking)60.4
12Step3 VL 10B60.3
13InternVL3-14B59.1
14Qwen 2.5 VL 7B Instruct55.8
15Qwen 2 VL 7B Instruct53.2

Interactive version: theaggregate.ai/benchmark?slug=sis-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.