OVO-S-Bench - Generative Spatial Reasoning: leaderboard

Metric: Level 3 (generative spatial reasoning: simulation, consistency checks, route planning) multiple-choice accuracy (%) on the 1,680-question evaluation set, prefix-access mode (128 frames sampled uniformly from the video before the query timestamp, question given with the frames), mean of its three task families. Source: arxiv.org. Saturation forecast: Around 2032. 31 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B61.2
2Qwen 3.5 397B A17B58.07
3Gemini 3.1 Pro (Preview)55.9
4Qwen 3 VL 4B Instruct54.54
5Gemini 3.1 Flash Lite54.14
6Qwen 3.5 27B (Non-reasoning)52.39
7Qwen 3 VL 32B Instruct51.94
8GPT-5.450.8
9Qwen 3.5 9B (Non-reasoning)50.2
10Qwen 3.5 4B (Non-reasoning)49.25
11Grok 4.1 Fast (Non-reasoning)48.49
12InternVL3.5-8B47.2
13Qwen 2.5 VL 7B Instruct45.86
14Gemma 4 26B A4B45.01
15Gemma 4 E4B42.8

Interactive version: theaggregate.ai/benchmark?slug=ovo-s-bench-generative-spatial-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.