Video-MME-Logical: leaderboard
Metric: Overall accuracy (%) over all test videos; Video-MME-Logical: 3,750 procedurally generated test videos in 25 task categories under five temporal-logical operations (state tracking, sequential counting, temporal ordering, dynamic spatiality, structural composition), each at easy, medium and hard settings; zero-shot, 2 fps; exact-match accuracy of the tagged multiple-choice or fill-in answer; human 95.9; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 28.6 |
| 2 | GPT-5.4 | 22.7 |
| 3 | Qwen 2.5 VL 72B Instruct | 12.5 |
| 4 | Qwen 3 VL 8B Instruct | 11.9 |
| 5 | Qwen 3 VL 30B A3B Instruct | 11.8 |
| 6 | Qwen 3 VL 30B A3B (Thinking) | 10.3 |
| 7 | Qwen 2.5 VL 7B Instruct | 7.4 |
| 8 | Qwen 3 VL 8B (Thinking) | 6.6 |
| 9 | Qwen3 Omni 30B A3B Instruct | 5.8 |
Interactive version: theaggregate.ai/benchmark?slug=video-mme-logical · How It Works · Data refreshed daily, snapshot 2026-09-29.