VSTAT: leaderboard

Metric: Average score (%) over all questions: 1,500 questions on 834 synthetic (Blender) and real-world video clips that require tracking a visual state through the whole video; accuracy on multiple-choice questions and mean relative accuracy on numerical questions (%), open models at their best of 16 to 128 uniformly sampled frames; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 23 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (Low)44.4
2Gemini 3.1 Pro (Preview) (High)43.9
3Gemini 3 Flash (High)38.8
4Molmo2-8B34
5Qwen 3 VL 8B Instruct33.2
6Qwen 3 VL 4B Instruct31.3
7InternVL3.5-8B30.6
8GLM-4.1V-9B (Thinking)30.2
9Qwen 3 VL 8B (Thinking)28.2
10Qwen 3 VL 4B (Thinking)26

Interactive version: theaggregate.ai/benchmark?slug=vstat · How It Works · Data refreshed daily, snapshot 2026-09-29.