ViMU - Evidence Grounding: leaderboard

Metric: Evidence grounding score (%): selecting the evidence sources (frames, on-screen text, editing, transcript, audio tone) that support the intended meaning, set-based multiple-choice scoring: 0 if any selected option is wrong, otherwise the share of gold options selected, averaged over items, on ViMU, short online videos whose meaning lies in subtext (irony, mockery, criticism), uniformly sampled frames, zero-shot through official implementations or APIs; questions and references written by GPT-5.4 and reviewed by human experts; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScore
1GPT-5.267.83
2Seed 2.0 Lite66.16
3GPT-5.4 Mini64.45
4Grok 4.1 Fast63.84
5Qwen 3.5 27B60.28
6Qwen 3 VL 32B Instruct59.64
7O4 Mini59.63
8Ministral 3 14B55.73
9Gemini 3 Flash (Preview)52.8
10Gemma 3 27B (IT)49.38
11MiMo-V2-Omni48.94
12Ministral 8B48.6
13Claude 3 Haiku34.55
14Gemma 3 4B (IT)25.41
15GLM-4.5V23.11

Interactive version: theaggregate.ai/benchmark?slug=vimu-evidence-grounding · How It Works · Data refreshed daily, snapshot 2026-10-07.