EYT-Bench (Nemotron-USA) - Emotion: leaderboard

Metric: Per-turn emotion accuracy (%; the target model first predicts the user emotion label from 15 classes, scored against the simulator gold label, before generating its reply; multi-turn dialogues of up to 10 turns with a persona-grounded Gemini 3.1 Pro user simulator and a Gemini 3.1 Pro judge (reasoning effort high, temperature 0); Nemotron-Personas-USA pool (100 dialogues; structured demographic personas)). Source: arxiv.org. Saturation forecast: Around August 2028. 17 models tracked.

Top models

#ModelScore
1GPT-5.550
2DeepSeek V4 Pro49.6
3Gemini 3.1 Pro (Preview) (Thinking)38.6
4Qwen 3.5 397B A17B33.7
5Claude Sonnet 4.632.6
6DeepSeek V4 Flash31.3
7Qwen 3.5 27B30.7
8Qwen 3.5 35B A3B27.2
9Claude Opus 4.725.3
10Seed 2.0 Pro24.6
11Seed 2.0 Mini23.3
12Seed 2.0 Lite20.3

Interactive version: theaggregate.ai/benchmark?slug=eyt-bench-nemotron-usa-emotion · How It Works · Data refreshed daily, snapshot 2026-09-29.