SIS-Bench - Positional Relationship: leaderboard

Metric: Accuracy (%; 252 positional relationship questions; four-option multiple choice on UAV videos (at most 32 frames), zero-shot, open models through vLLM with at most 128 generated tokens and proprietary models through their APIs; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 26 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)89.7
2Qwen 3.5 Plus87.7
3Seed 1.887.7
4Kimi K2.587.3
5GPT-5.486.9
6Qwen 3 VL 30B A3B Instruct85.3
7Qwen 3 VL 8B Instruct82.9
8Qwen 3 VL 4B Instruct82.1
9InternVL3-14B81.3
10Qwen 3 VL 8B (Thinking)79.8
11GLM-4.1V-9B (Thinking)75.4
12InternVL3.5-8B73.8
13Qwen 2.5 VL 7B Instruct69
14Qwen 2 VL 7B Instruct69
15Step3 VL 10B67.5

Interactive version: theaggregate.ai/benchmark?slug=sis-bench-positional-relationship · How It Works · Data refreshed daily, snapshot 2026-09-29.