ViMU: leaderboard

Metric: Average (%) of the four task scores (open-ended interpretation, evidence grounding, rhetoric mechanism and social value signal identification) on ViMU, short online videos whose meaning lies in subtext (irony, mockery, criticism), uniformly sampled frames, zero-shot through official implementations or APIs; questions and references written by GPT-5.4 and reviewed by human experts; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 16 models tracked.

Top models

#ModelScore
1O4 Mini46.91
2Grok 4.1 Fast46.28
3Qwen 3.5 27B45.91
4GPT-5.244.67
5Gemini 3 Flash (Preview)44.31
6Qwen 3 VL 32B Instruct41.64
7Seed 2.0 Lite40.62
8MiMo-V2-Omni38.14
9GPT-5.4 Mini36.64
10Gemma 3 27B (IT)36.43
11Ministral 3 14B35.45
12Ministral 8B34.79
13GLM-4.5V25.94
14Gemma 3 4B (IT)23.28
15Claude 3 Haiku22.9

Interactive version: theaggregate.ai/benchmark?slug=vimu · How It Works · Data refreshed daily, snapshot 2026-10-07.