EIBench (Emotion Management) - DeepSeek-V4-Pro Simulator: leaderboard

Metric: Anchor-based outcome score x 100 (-100 to +100): the simulated user's final emotion and relation state mapped per axis to 0 at the start anchor, +1 at the success anchor and -1 at the failure anchor (clipped), scenario-weighted average over the four scene types; 213 held-out test scenarios, multi-turn dialogue with an LLM user simulator, DeepSeek-V4-Pro as the simulator; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.624
2DeepSeek V4 Pro23.3
3GPT-5.422.7
4Kimi K2.621.9
5Gemini 3.1 Pro (Preview)21.5
6Qwen 3.6 Max Preview19.9
7GLM-5.119.4
8MiniMax-M2.517.6
9Seed 2.0 Pro17
10Gemini 3 Flash16.5
11Qwen 3 Max14.8
12Grok 4.2013
13Grok 4.20 (Reasoning)10.6
14Qwen 3 32B5.8
15Qwen 3 8B-23.7

Interactive version: theaggregate.ai/benchmark?slug=eibench-emotion-management-deepseek-v4-pro-simulator · How It Works · Data refreshed daily, snapshot 2026-09-29.