LongVideoBench — leaderboard

Long-form video understanding benchmark requiring models to reason over extended temporal contexts across diverse video content.

Metric: Test Total Accuracy (%). Source: longvideobench.github.io. Status: saturation imminent. 35 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro64.4
2Gemini 1.5 Flash62.4
3GPT-4 Turbo60.7
4GPT-4o Mini58.8
5MiniCPM-V-2.657.7
6Qwen 2 VL 7B56.8

Interactive version: theaggregate.ai/benchmark?slug=longvideobench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.