DAR (Dynamic Affective Reasoning) - Reasoning Quality: leaderboard

Metric: GPT-Score (out of 5): GPT-4o judge rating of each rationale for visual grounding, causal logic, viewer centricity, temporal consistency and answer consistency, averaged over the five; 1,441 test videos of the DAR viewer-centric emotion dataset (27 emotion categories, emotions changing across consecutive events); the model segments the video into affective segments, labels each segment's emotion and writes a causal rationale; off-the-shelf models without fine-tuning on DAR. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1Qwen 2.5 VL 7B Instruct2.6
2InternVL3.5-8B2.2

Interactive version: theaggregate.ai/benchmark?slug=dar-dynamic-affective-reasoning-reasoning-quality · How It Works · Data refreshed daily, snapshot 2026-09-29.