E3mo-Bench (Expressed) - Open-Vocabulary Recognition (VAD Similarity): leaderboard

Metric: Redundancy-aware VAD set similarity F-VAD (%; predicted and reference open-vocabulary emotion sets compared in a z-normalized valence-arousal-dominance lexicon space after merging near-duplicate emotions, so that distant confusions cost more; 16 MLLMs on videos from eight emotion datasets, each assigned one affective perspective; subset: the emotion expressed by the people on screen). Source: arxiv.org. Saturation forecast: Around 2033. 16 models tracked.

Top models

#ModelScore
1GPT-5.448.89
2InternVL3-8B48.41
3Qwen2.5-Omni-7B44.47
4Claude Sonnet 4.643.78

Interactive version: theaggregate.ai/benchmark?slug=e3mo-bench-expressed-open-vocabulary-recognition-vad-similarity · How It Works · Data refreshed daily, snapshot 2026-09-29.