ViMU - Structured Subtext Understanding: leaderboard

Metric: Structured subtext understanding (%): the mean of the rhetoric mechanism and social value signal identification scores on ViMU, short online videos whose meaning lies in subtext (irony, mockery, criticism), uniformly sampled frames, zero-shot through official implementations or APIs; questions and references written by GPT-5.4 and reviewed by human experts; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 16 models tracked.

Top models

#ModelScore
1Grok 4.1 Fast31.82
2O4 Mini31.36
3Gemini 3 Flash (Preview)30.94
4Qwen 3.5 27B30.29
5Qwen 3 VL 32B Instruct21.41
6Ministral 8B21.16
7Gemma 3 27B (IT)20.21
8MiMo-V2-Omni19.78
9GPT-5.218.85
10Seed 2.0 Lite17.74
11Ministral 3 14B16.93
12Gemma 3 4B (IT)14.13
13GLM-4.5V9.06
14GPT-5.4 Mini7.97
15GPT-4.1 Nano5.67

Interactive version: theaggregate.ai/benchmark?slug=vimu-structured-subtext-understanding · How It Works · Data refreshed daily, snapshot 2026-10-07.