VRBench: leaderboard

Long narrative video benchmark for multi-step reasoning, assessing both final outcomes and process-level reasoning chains over temporally grounded questions.

Metric: Overall (self-reported). Source: benchmarklist.com. Status: saturated. 29 models tracked.

Top models

#ModelScore
1GPT-4o68.68
2Claude 3.7 Sonnet68.15
3InternVL2.5-78B62.31
4Gemini 2.0 Flash (Thinking)60.97
5O1 Preview60.14
6DeepSeek R157.13
7DeepSeek V356.06
8Qwen 2 VL 7B54.08
9Qwen 2.5 72B Instruct53.51
10QwQ-32B52.52
11Llama 3.3 70B Instruct49.84
12Qwen 2.5 7B Instruct48.29
13QwQ 32B-Preview35.9

Interactive version: theaggregate.ai/benchmark?slug=vrbench · How It Works · Data refreshed daily, snapshot 2026-09-05.