E3mo-Bench (Evoked) - Open-Vocabulary Recognition (VAD Similarity): leaderboard

Metric: Redundancy-aware VAD set similarity F-VAD (%; predicted and reference open-vocabulary emotion sets compared in a z-normalized valence-arousal-dominance lexicon space after merging near-duplicate emotions, so that distant confusions cost more; 16 MLLMs on videos from eight emotion datasets, each assigned one affective perspective; subset: the viewer's evoked emotion). Source: arxiv.org. Saturation forecast: Around 2032. 16 models tracked.

Top models

#ModelScore
1GPT-5.455.38
2InternVL3-8B50.54
3Qwen2.5-Omni-7B49.36
4Claude Sonnet 4.648.76

Interactive version: theaggregate.ai/benchmark?slug=e3mo-bench-evoked-open-vocabulary-recognition-vad-similarity · How It Works · Data refreshed daily, snapshot 2026-09-29.