MultivationBench - Example F1: leaderboard
Metric: Example-based F1 (%, partial credit per question) of multi-label motivation predictions for visually grounded character behaviors in sequential visual narratives, reasoning over the accumulated story images and text (Maslow 8-level needs and Reiss 16 basic desires, definition and practical-motivation tasks), zero-shot at temperature 0; multimodal input, common subset of 14,180 tasks from 1,000 stories that every model answered; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.1 Fast | 55.32 |
| 2 | Gemini 3 Flash | 54.22 |
| 3 | O4 Mini | 53.62 |
| 4 | Llama 4 Maverick | 49.81 |
| 5 | Llama 4 Scout | 49.6 |
| 6 | Phi-4 Multimodal Instruct | 43.03 |
| 7 | Nemotron Nano 12B v2 VL | 38.16 |
Interactive version: theaggregate.ai/benchmark?slug=multivationbench-example-f1 · How It Works · Data refreshed daily, snapshot 2026-09-29.