SIS-Bench - Action Sequence: leaderboard

Metric: Accuracy (%; 315 action sequence questions; four-option multiple choice on UAV videos (at most 32 frames), zero-shot, open models through vLLM with at most 128 generated tokens and proprietary models through their APIs; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 26 models tracked.

Top models

#ModelScore
1Seed 1.885.1
2Gemini 3 Flash (Preview)84.1
3Qwen 3.5 Plus82.5
4GPT-5.477.8
5Kimi K2.577.8
6Qwen 3 VL 8B (Thinking)62.9
7Qwen 3 VL 4B Instruct61.3
8Qwen 3 VL 8B Instruct60.3
9GLM-4.1V-9B (Thinking)55.9
10Step3 VL 10B55.9
11Qwen 2.5 VL 7B Instruct55.6
12InternVL3.5-8B52.4
13Qwen 2 VL 7B Instruct51.4
14Qwen 3 VL 30B A3B Instruct44.8
15InternVL3-14B43.8

Interactive version: theaggregate.ai/benchmark?slug=sis-bench-action-sequence · How It Works · Data refreshed daily, snapshot 2026-09-29.