MLVU — leaderboard

Multi-task Long Video Understanding benchmark with diverse evaluation tasks across multiple video lengths and genres.

Metric: M-Avg (%). Source: github.com. Status: saturated. 41 models tracked.

Top models

#ModelScore
1Full mark100
2[LVAgent](https://github.com/64327069/LVAgent)83.9
3[TSPO](https://github.com/Hui-design/TSPO)77.3
4[TSPO](https://github.com/Hui-design/TSPO)76.3
5[VideoChat-Flash](https://github.com/OpenGVLab/VideoChat-Flash)74.7
6[VideoLLaMA3](https://github.com/DAMO-NLP-SG/VideoLLaMA3)73
7[Oryx-1.5](https://oryx-mllm.github.io)72.3
8[Aria](https://rhymes.ai/blog-details/aria-first-open-multimodal-native-moe-model)70.6
9[LinVT](https://github.com/gls0425/LinVT)68.9
10[LLaVA-OneVision](https://github.com/LLaVA-VL/LLaVA-NeXT/tree/main)66.4

Interactive version: theaggregate.ai/benchmark?slug=mlvu · How the rankings work · Data refreshed daily, snapshot 2026-07-22.