MVPBench - Caption-Free Temporal Splicing: leaderboard

Metric: MVPBench normalized proficiency score (accuracy minus chance, divided by one minus chance, times 100, so 0 is chance level and 100 is perfect) on its 89 multi-video questions ordering shuffled clips of a process with inherent temporal logic from the visuals alone; chance is one over the number of possible clip orders, which varies by item; answers extracted by rules and GPT-4-turbo; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash11.54#237
2GPT-4o10.21#333
3Qwen 2.5 VL 72B Instruct1.45#364
4InternVL3-78B-2.92#345
5Qwen 2.5 VL 7B Instruct-5.12#643
6InternVL3-38B-7.3#395
7InternVL3-8B-11.92#606

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mvpbench-caption-free-temporal-splicing · How It Works · Data refreshed daily, snapshot 2026-10-11.