AutoEval-Video — leaderboard

AutoEval-Video evaluates large vision-language models on open-ended video question answering across nine video perception and reasoning dimensions.

Metric: Overall Accuracy (%). Source: huggingface.co. 8 models tracked.

Top models

#ModelScore
1GPT-4V22.2
2VideoChat13.4
3Video-LLaMA11.2
4LLaVA-1.58.5
5InstructBLIP7.9
6Qwen-VL7
7Video-ChatGPT6.7
8BLIP-20.3

Interactive version: theaggregate.ai/benchmark?slug=autoeval-video · How the rankings work · Data refreshed daily, snapshot 2026-07-22.