EYT-Bench (PersonaMem-v2) - Empathy: leaderboard

Metric: Empathy (recognition, attribution, resonance, response, support) (0-100): judge rubric of five sub-indicators scored 0, 0.5 or 1 per turn, summed to 0-5, turn weighted with the first two turns at 0.10 each and rescaled to 0-100; multi-turn dialogues of up to 10 turns with a persona-grounded Gemini 3.1 Pro user simulator and a Gemini 3.1 Pro judge (reasoning effort high, temperature 0); PersonaMem-v2 pool (100 dialogues; personas are free-text paragraphs distilled from long real user-assistant interactions). Source: arxiv.org. Saturation forecast: Around 2030. 17 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B73.1
2Qwen 3.5 27B71.3
3Qwen 3.5 35B A3B70.1
4Seed 2.0 Pro69.4
5DeepSeek V4 Pro67.8
6Seed 2.0 Lite67.8
7Gemini 3.1 Pro (Preview) (Thinking)67.8
8DeepSeek V4 Flash67.6
9Claude Opus 4.765.3
10Claude Sonnet 4.662.9
11Seed 2.0 Mini57.6
12GPT-5.555.9

Interactive version: theaggregate.ai/benchmark?slug=eyt-bench-personamem-v2-empathy · How It Works · Data refreshed daily, snapshot 2026-09-29.