DAR (Dynamic Affective Reasoning) - Reasoning Quality: leaderboard
Metric: GPT-Score (out of 5): GPT-4o judge rating of each rationale for visual grounding, causal logic, viewer centricity, temporal consistency and answer consistency, averaged over the five; 1,441 test videos of the DAR viewer-centric emotion dataset (27 emotion categories, emotions changing across consecutive events); the model segments the video into affective segments, labels each segment's emotion and writes a causal rationale; off-the-shelf models without fine-tuning on DAR. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 VL 7B Instruct | 2.6 |
| 2 | InternVL3.5-8B | 2.2 |
Interactive version: theaggregate.ai/benchmark?slug=dar-dynamic-affective-reasoning-reasoning-quality · How It Works · Data refreshed daily, snapshot 2026-09-29.