MultivationBench (Text-Only): leaderboard
Metric: Exact-match accuracy (%) of multi-label motivation predictions for visually grounded character behaviors in sequential visual narratives, reasoning over the accumulated story images and text (Maslow 8-level needs and Reiss 16 basic desires, definition and practical-motivation tasks), zero-shot at temperature 0; caption-enriched text-only input instead of images, all four task types and all story lengths; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.1 Fast | 30.6 |
| 2 | Llama 4 Scout | 26.7 |
| 3 | O4 Mini | 23.5 |
| 4 | Gemini 3 Flash | 13.6 |
| 5 | Llama 4 Maverick | 8.5 |
| 6 | Nemotron Nano 12B v2 VL | 8.5 |
| 7 | Phi-4 Multimodal Instruct | 5.9 |
Interactive version: theaggregate.ai/benchmark?slug=multivationbench-text-only · How It Works · Data refreshed daily, snapshot 2026-09-29.