MultivationBench - Example F1: leaderboard

Metric: Example-based F1 (%, partial credit per question) of multi-label motivation predictions for visually grounded character behaviors in sequential visual narratives, reasoning over the accumulated story images and text (Maslow 8-level needs and Reiss 16 basic desires, definition and practical-motivation tasks), zero-shot at temperature 0; multimodal input, common subset of 14,180 tasks from 1,000 stories that every model answered; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.

Top models

#ModelScore
1Grok 4.1 Fast55.32
2Gemini 3 Flash54.22
3O4 Mini53.62
4Llama 4 Maverick49.81
5Llama 4 Scout49.6
6Phi-4 Multimodal Instruct43.03
7Nemotron Nano 12B v2 VL38.16

Interactive version: theaggregate.ai/benchmark?slug=multivationbench-example-f1 · How It Works · Data refreshed daily, snapshot 2026-09-29.