ViMU - Rhetoric Mechanism Identification: leaderboard
Metric: Rhetoric mechanism identification score (%): selecting the rhetorical devices (irony, exaggeration, juxtaposition and others) that carry the meaning, set-based multiple-choice scoring: 0 if any selected option is wrong, otherwise the share of gold options selected, averaged over items, on ViMU, short online videos whose meaning lies in subtext (irony, mockery, criticism), uniformly sampled frames, zero-shot through official implementations or APIs; questions and references written by GPT-5.4 and reviewed by human experts; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 27B | 38.18 |
| 2 | Grok 4.1 Fast | 34.91 |
| 3 | Gemini 3 Flash (Preview) | 33.63 |
| 4 | O4 Mini | 33.21 |
| 5 | Gemma 3 27B (IT) | 32.47 |
| 6 | Ministral 8B | 31.87 |
| 7 | Qwen 3 VL 32B Instruct | 27.65 |
| 8 | Ministral 3 14B | 27.29 |
| 9 | Gemma 3 4B (IT) | 21.1 |
| 10 | MiMo-V2-Omni | 21.04 |
| 11 | Seed 2.0 Lite | 18.75 |
| 12 | GPT-5.2 | 16.55 |
| 13 | GLM-4.5V | 8.87 |
| 14 | GPT-5.4 Mini | 4.17 |
| 15 | Claude 3 Haiku | 2.99 |
Interactive version: theaggregate.ai/benchmark?slug=vimu-rhetoric-mechanism-identification · How It Works · Data refreshed daily, snapshot 2026-10-07.