EmoTrans - Emotion Transition Reasoning: leaderboard

Metric: LLM-Score (%) on Emotion Transition Reasoning (724 questions: a GPT-5 judge checks whether the model's explanation of why the emotion changed or stayed matches the reference explanation), on EmoTrans, built from 1,000 manually annotated real-world video clips (OpenHumanVid and CH-SIMS v2) over 12 scenarios with one or more target persons; zero-shot, video with its text (and audio where the model accepts it); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1GPT-5.485.7
2Gemini 2.5 Flash (Non-reasoning)82.7
3Gemini 3 Pro82.3
4Seed 2.0 Lite81.4
5GLM-4.5V79.5
6GLM-4.6V79
7Qwen 3 VL 32B Instruct78.9
8Qwen 3.5 27B77.9
9Qwen 3.5 122B A10B77.3
10Qwen 2.5 VL 72B Instruct76.4

Interactive version: theaggregate.ai/benchmark?slug=emotrans-emotion-transition-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.