OpenVLM Video - Video-MME (w/o subs) - Spatial Perception: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 37 models tracked.

Top models

#ModelScore
1GPT-4o (2024-08-06)79.6
2Gemini 2.0 Flash77.8
3InternVL3-78B77.8
4InternVL2.5-78B74.1
5Qwen 2 VL 72B70.4
6Qwen 2 VL 7B66.7
7InternVL3-8B64.8
8InternVL2-8B63
9Gemini 1.5 Flash63
10Claude 3.5 Sonnet59.3
11Phi-4 Multimodal Instruct46.3

Interactive version: theaggregate.ai/benchmark?slug=openvlm-video-video-mme-w-o-subs-spatial-perception · How It Works · Data refreshed daily, snapshot 2026-09-19.