SIS-Bench - Action Prediction: leaderboard

Metric: Accuracy (%; 263 action prediction questions; four-option multiple choice on UAV videos (at most 32 frames), zero-shot, open models through vLLM with at most 128 generated tokens and proprietary models through their APIs; higher is better). Source: arxiv.org. Saturation forecast: Around 2029. 26 models tracked.

Top models

#ModelScore
1Qwen 3.5 Plus69.2
2GPT-5.465.8
3Kimi K2.565.4
4Qwen 3 VL 8B (Thinking)64.6
5Seed 1.863.1
6Gemini 3 Flash (Preview)61.6
7Qwen 2.5 VL 7B Instruct60.8
8Qwen 3 VL 30B A3B Instruct59.7
9Qwen 2 VL 7B Instruct58.9
10Qwen 3 VL 8B Instruct57
11Qwen 3 VL 4B Instruct57
12InternVL3-14B55.5
13InternVL3.5-8B55.1
14Step3 VL 10B54.4
15GLM-4.1V-9B (Thinking)48.7

Interactive version: theaggregate.ai/benchmark?slug=sis-bench-action-prediction · How It Works · Data refreshed daily, snapshot 2026-09-29.