EmoTrans - Emotion Transition Reasoning: leaderboard
Metric: LLM-Score (%) on Emotion Transition Reasoning (724 questions: a GPT-5 judge checks whether the model's explanation of why the emotion changed or stayed matches the reference explanation), on EmoTrans, built from 1,000 manually annotated real-world video clips (OpenHumanVid and CH-SIMS v2) over 12 scenarios with one or more target persons; zero-shot, video with its text (and audio where the model accepts it); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 85.7 |
| 2 | Gemini 2.5 Flash (Non-reasoning) | 82.7 |
| 3 | Gemini 3 Pro | 82.3 |
| 4 | Seed 2.0 Lite | 81.4 |
| 5 | GLM-4.5V | 79.5 |
| 6 | GLM-4.6V | 79 |
| 7 | Qwen 3 VL 32B Instruct | 78.9 |
| 8 | Qwen 3.5 27B | 77.9 |
| 9 | Qwen 3.5 122B A10B | 77.3 |
| 10 | Qwen 2.5 VL 72B Instruct | 76.4 |
Interactive version: theaggregate.ai/benchmark?slug=emotrans-emotion-transition-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.