EYT-Bench (Nemotron-USA) - Empathy: leaderboard

Metric: Empathy (recognition, attribution, resonance, response, support) (0-100): judge rubric of five sub-indicators scored 0, 0.5 or 1 per turn, summed to 0-5, turn weighted with the first two turns at 0.10 each and rescaled to 0-100; multi-turn dialogues of up to 10 turns with a persona-grounded Gemini 3.1 Pro user simulator and a Gemini 3.1 Pro judge (reasoning effort high, temperature 0); Nemotron-Personas-USA pool (100 dialogues; structured demographic personas). Source: arxiv.org. Saturation forecast: Around 2029. 17 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B77.4
2Qwen 3.5 35B A3B76.7
3Qwen 3.5 397B A17B76.3
4Gemini 3.1 Pro (Preview) (Thinking)73.8
5Claude Sonnet 4.673.4
6Claude Opus 4.772.3
7DeepSeek V4 Pro70.7
8Seed 2.0 Lite70.1
9DeepSeek V4 Flash69.8
10Seed 2.0 Pro69
11Seed 2.0 Mini63.9
12GPT-5.557

Interactive version: theaggregate.ai/benchmark?slug=eyt-bench-nemotron-usa-empathy · How It Works · Data refreshed daily, snapshot 2026-09-29.