VRBench — leaderboard

Long narrative video benchmark for multi-step reasoning, assessing both final outcomes and process-level reasoning chains over temporally grounded questions.

Metric: Overall (self-reported). Source: benchmarklist.com. Status: saturation imminent. 29 models tracked.

Top models

#ModelScore
1GPT-4o68.68
2Claude 3.7 Sonnet68.15
3InternVL2.5-78B62.31
4Gemini 2.0 Flash (Thinking)60.97
5O1 Preview60.14
6DeepSeek V356.06
7Qwen 2 VL 7B54.08
8Qwen 2.5 72B Instruct53.51
9QwQ-32B52.52
10Llama 3.3 70B Instruct49.84
11Qwen 2.5 7B Instruct48.29
12QwQ 32B-Preview35.9

Interactive version: theaggregate.ai/benchmark?slug=vrbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.