MultivationBench: leaderboard

Metric: Exact-match accuracy (%, predicted option set equals the gold set) of multi-label motivation predictions for visually grounded character behaviors in sequential visual narratives, reasoning over the accumulated story images and text (Maslow 8-level needs and Reiss 16 basic desires, definition and practical-motivation tasks), zero-shot at temperature 0; multimodal input, common subset of 14,180 tasks from 1,000 stories that every model answered; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 8 models tracked.

Top models

#ModelScore
1Phi-4 Multimodal Instruct39.22
2Gemini 3 Flash36.09
3Llama 4 Scout33.48
4Grok 4.1 Fast32.53
5Llama 4 Maverick30.28
6O4 Mini27.29
7Nemotron Nano 12B v2 VL8.58

Interactive version: theaggregate.ai/benchmark?slug=multivationbench · How It Works · Data refreshed daily, snapshot 2026-09-29.