VGenST-Bench - Height Ordering: leaderboard

Metric: Height Ordering (vista scale, exocentric, static) accuracy (%) on the base multiple-choice questions of VGenST-Bench under circular evaluation (correct only if right under every cyclic permutation of the options), macro-averaged over the task's QA types, 8 uniformly sampled frames from about 100 synthesized videos per task (generated with text-to-image and image-to-video models from validated scene graphs); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Flash96.9
2Gemini 3.1 Flash Lite87.1
3Kimi K2.686.1
4GPT-5.484.5
5GPT-5.4 Mini70.7
6Qwen 3.5 27B69.5
7Qwen 3.5 9B67.2
8Gemma 4 31B (IT)66.8
9Qwen 3.5 4B63.6
10Gemma 4 26B A4B (IT)61.9
11GPT-5.4 Nano44

Interactive version: theaggregate.ai/benchmark?slug=vgenst-bench-height-ordering · How It Works · Data refreshed daily, snapshot 2026-10-07.