LVBench: leaderboard

1,549 four-option questions on 103 videos averaging 68 minutes across six skills (temporal grounding, summarization, entity recognition), from Zhipu AI and Tsinghua (2024); humans 94.4%.

Metric: Score (self-reported). Source: benchmarklist.com. Status: years away from saturation. 43 models tracked.

Top models

#ModelScore
1Gemini 3.8 Flash87.8
2Gemini 3.7 Flash85.4
3Gemini 3.6 Flash84.2
4GPT-5.6 Sol82.1
5Qwen 3.8 Max81.8
6GPT-5.6 Terra78.9
7Seed 2.1 Pro78
8GPT-5.4 (xHigh)77.4
9Seed 2.1 Turbo76.8
10Qwen3.8-Flash-Next76.6
11Qwen 3.8 Flash76.6
12Qwen 3.7 Plus76.2
13Kimi K2.575.9
14Claude Opus 575.4
15Gemini 3.1 Pro (Preview)75.1

Interactive version: theaggregate.ai/benchmark?slug=lvbench · How It Works · Data refreshed daily, snapshot 2026-09-05.