ViMU - Rhetoric Mechanism Identification: leaderboard

Metric: Rhetoric mechanism identification score (%): selecting the rhetorical devices (irony, exaggeration, juxtaposition and others) that carry the meaning, set-based multiple-choice scoring: 0 if any selected option is wrong, otherwise the share of gold options selected, averaged over items, on ViMU, short online videos whose meaning lies in subtext (irony, mockery, criticism), uniformly sampled frames, zero-shot through official implementations or APIs; questions and references written by GPT-5.4 and reviewed by human experts; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 16 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B38.18
2Grok 4.1 Fast34.91
3Gemini 3 Flash (Preview)33.63
4O4 Mini33.21
5Gemma 3 27B (IT)32.47
6Ministral 8B31.87
7Qwen 3 VL 32B Instruct27.65
8Ministral 3 14B27.29
9Gemma 3 4B (IT)21.1
10MiMo-V2-Omni21.04
11Seed 2.0 Lite18.75
12GPT-5.216.55
13GLM-4.5V8.87
14GPT-5.4 Mini4.17
15Claude 3 Haiku2.99

Interactive version: theaggregate.ai/benchmark?slug=vimu-rhetoric-mechanism-identification · How It Works · Data refreshed daily, snapshot 2026-10-07.