VidMsg QA: leaderboard
Metric: Overall accuracy (%), five-option multiple choice: pick the intended implicit message of a short YouTube clip from distractors sampled within the same topic (one question per clip over 9 topics and 52 target messages; the per-topic cells imply 395 scored clips); chance 20; visual frames only (native frame protocol of each model, at most 1 frame per second); higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 VL 7B Instruct | 71.9 |
| 2 | GPT-5.4 Mini | 70.6 |
| 3 | GPT-5.4 (2026-03-05) | 70.4 |
| 4 | Qwen 3 VL 32B | 68.9 |
| 5 | Qwen 2.5 VL 32B Instruct | 67.6 |
| 6 | Molmo2-8B | 57.7 |
| 7 | MiniCPM-V-4.5-8B | 22.3 |
Interactive version: theaggregate.ai/benchmark?slug=vidmsg-qa · How It Works · Data refreshed daily, snapshot 2026-09-29.