MultivationBench (Text-Only): leaderboard

Metric: Exact-match accuracy (%) of multi-label motivation predictions for visually grounded character behaviors in sequential visual narratives, reasoning over the accumulated story images and text (Maslow 8-level needs and Reiss 16 basic desires, definition and practical-motivation tasks), zero-shot at temperature 0; caption-enriched text-only input instead of images, all four task types and all story lengths; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.

Top models

#ModelScore
1Grok 4.1 Fast30.6
2Llama 4 Scout26.7
3O4 Mini23.5
4Gemini 3 Flash13.6
5Llama 4 Maverick8.5
6Nemotron Nano 12B v2 VL8.5
7Phi-4 Multimodal Instruct5.9

Interactive version: theaggregate.ai/benchmark?slug=multivationbench-text-only · How It Works · Data refreshed daily, snapshot 2026-09-29.