VSTAT - Dictionary: leaderboard

Metric: Score (%) on questions about dictionary (entity-to-value) state structures: 1,500 questions on 834 synthetic (Blender) and real-world video clips that require tracking a visual state through the whole video; accuracy on multiple-choice questions and mean relative accuracy on numerical questions (%), open models at their best of 16 to 128 uniformly sampled frames; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 23 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (High)39.3
2Gemini 3.1 Pro (Preview) (Low)38.7
3Gemini 3 Flash (High)32.4
4Qwen 3 VL 8B Instruct31.5
5InternVL3.5-8B28.3
6Qwen 3 VL 8B (Thinking)28
7GLM-4.1V-9B (Thinking)27.3
8Molmo2-8B27
9Qwen 3 VL 4B Instruct25.8
10Qwen 3 VL 4B (Thinking)21.4

Interactive version: theaggregate.ai/benchmark?slug=vstat-dictionary · How It Works · Data refreshed daily, snapshot 2026-09-29.