ViMU - Open-Ended Interpretation: leaderboard

Metric: Open-ended interpretation score (%): an LLM judge rubric over core intent, implicit signal and social target, minus hallucination and literal-only penalties, rescaled to percent, hint-free questions on ViMU, short online videos whose meaning lies in subtext (irony, mockery, criticism), uniformly sampled frames, zero-shot through official implementations or APIs; questions and references written by GPT-5.4 and reviewed by human experts; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 16 models tracked.

Top models

#ModelScore
1GPT-5.273.15
2GPT-5.4 Mini66.19
3O4 Mini65.27
4Qwen 3 VL 32B Instruct64.09
5MiMo-V2-Omni64.07
6Qwen 3.5 27B62.8
7Gemini 3 Flash (Preview)62.54
8GLM-4.5V62.52
9Seed 2.0 Lite60.84
10Grok 4.1 Fast57.62
11Gemma 3 27B (IT)55.9
12Ministral 3 14B52.19
13Claude 3 Haiku50.41
14GPT-4.1 Nano50.12
15Ministral 8B48.25

Interactive version: theaggregate.ai/benchmark?slug=vimu-open-ended-interpretation · How It Works · Data refreshed daily, snapshot 2026-10-07.