SIS-Bench - Action Recognition: leaderboard

Metric: Accuracy (%; 686 action recognition questions; four-option multiple choice on UAV videos (at most 32 frames), zero-shot, open models through vLLM with at most 128 generated tokens and proprietary models through their APIs; higher is better). Source: arxiv.org. Saturation forecast: Around 2030. 26 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)66.5
2GPT-5.457.3
3InternVL3.5-8B56.1
4Qwen 3 VL 8B Instruct55.1
5Qwen 3 VL 4B Instruct53.4
6Qwen 3 VL 8B (Thinking)53.4
7Qwen 3.5 Plus52.8
8Kimi K2.550.9
9Step3 VL 10B50
10GLM-4.1V-9B (Thinking)48.1
11Seed 1.843.7
12Qwen 2.5 VL 7B Instruct43.3
13InternVL3-14B42.7
14Qwen 3 VL 30B A3B Instruct39.7
15Qwen 2 VL 7B Instruct38.6

Interactive version: theaggregate.ai/benchmark?slug=sis-bench-action-recognition · How It Works · Data refreshed daily, snapshot 2026-09-29.