VSTAT - Set: leaderboard

Metric: Score (%) on questions about set state structures: 1,500 questions on 834 synthetic (Blender) and real-world video clips that require tracking a visual state through the whole video; accuracy on multiple-choice questions and mean relative accuracy on numerical questions (%), open models at their best of 16 to 128 uniformly sampled frames; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 23 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (Low)51.9
2Gemini 3.1 Pro (Preview) (High)50
3Gemini 3 Flash (High)48.4
4InternVL3.5-8B41.8
5GLM-4.1V-9B (Thinking)40.8
6Qwen 3 VL 4B Instruct39.8
7Molmo2-8B39.1
8Qwen 3 VL 8B Instruct37.9
9Qwen 3 VL 4B (Thinking)29.5
10Qwen 3 VL 8B (Thinking)28.5

Interactive version: theaggregate.ai/benchmark?slug=vstat-set · How It Works · Data refreshed daily, snapshot 2026-09-29.