ViSTR-Bench: leaderboard

Metric: Accuracy (%; averaged over all 1,340 two-option questions in 15 video subtasks across motion perception, spatial relations, outcome prediction and physical dynamics; chance 50, majority-class 57.9). Source: arxiv.org. Saturation forecast: Around 2032. 39 models tracked.

Top models

#ModelScore
1GPT-5.4 (Thinking)62
2Seed 2.0 Pro (Thinking)60.1
3Seed 2.0 Pro (Non-reasoning)56.8
4GPT-5.4 (Non-reasoning)56.1
5Claude Opus 4.655.4
6Qwen 3.5 27B55
7Claude Opus 4.6 (Thinking)54.7
8Qwen 3.5 397B A17B54.4
9Gemini 3.1 Pro (Preview)53.6
10Qwen 3.5 122B A10B53.4
11Claude Sonnet 4.6 (Thinking)53.4
12Gemini 3.1 Flash Lite (Preview)52.6
13Intern-S151.7
14Claude Sonnet 4.651
15InternVL3.5-8B50.9

Interactive version: theaggregate.ai/benchmark?slug=vistr-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.