Video-MME-Logical: leaderboard

Metric: Overall accuracy (%) over all test videos; Video-MME-Logical: 3,750 procedurally generated test videos in 25 task categories under five temporal-logical operations (state tracking, sequential counting, temporal ordering, dynamic spatiality, structural composition), each at easy, medium and hard settings; zero-shot, 2 fps; exact-match accuracy of the tagged multiple-choice or fill-in answer; human 95.9; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 17 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)28.6
2GPT-5.422.7
3Qwen 2.5 VL 72B Instruct12.5
4Qwen 3 VL 8B Instruct11.9
5Qwen 3 VL 30B A3B Instruct11.8
6Qwen 3 VL 30B A3B (Thinking)10.3
7Qwen 2.5 VL 7B Instruct7.4
8Qwen 3 VL 8B (Thinking)6.6
9Qwen3 Omni 30B A3B Instruct5.8

Interactive version: theaggregate.ai/benchmark?slug=video-mme-logical · How It Works · Data refreshed daily, snapshot 2026-09-29.