LVBench: leaderboard
1,549 four-option questions on 103 videos averaging 68 minutes across six skills (temporal grounding, summarization, entity recognition), from Zhipu AI and Tsinghua (2024); humans 94.4%.
Metric: Score (self-reported). Source: benchmarklist.com. Status: years away from saturation. 43 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.8 Flash | 87.8 |
| 2 | Gemini 3.7 Flash | 85.4 |
| 3 | Gemini 3.6 Flash | 84.2 |
| 4 | GPT-5.6 Sol | 82.1 |
| 5 | Qwen 3.8 Max | 81.8 |
| 6 | GPT-5.6 Terra | 78.9 |
| 7 | Seed 2.1 Pro | 78 |
| 8 | GPT-5.4 (xHigh) | 77.4 |
| 9 | Seed 2.1 Turbo | 76.8 |
| 10 | Qwen3.8-Flash-Next | 76.6 |
| 11 | Qwen 3.8 Flash | 76.6 |
| 12 | Qwen 3.7 Plus | 76.2 |
| 13 | Kimi K2.5 | 75.9 |
| 14 | Claude Opus 5 | 75.4 |
| 15 | Gemini 3.1 Pro (Preview) | 75.1 |
Interactive version: theaggregate.ai/benchmark?slug=lvbench · How It Works · Data refreshed daily, snapshot 2026-09-05.