EIBench (Emotion Management) - Support: leaderboard

Metric: Anchor-based outcome score x 100 (-100 to +100): the simulated user's final emotion and relation state mapped per axis to 0 at the start anchor, +1 at the success anchor and -1 at the failure anchor (clipped), averaged over the 51 Support scenarios (supporting a distressed user); 213 held-out test scenarios, multi-turn dialogue with an LLM user simulator, Qwen3-Max as the simulator; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.656.5
2GPT-5.449
3Qwen 3 Max45.7
4GLM-5.144.8
5Qwen 3.6 Max Preview42.7
6Kimi K2.642.1
7Grok 4.20 (Reasoning)40.9
8Seed 2.0 Pro39.8
9Gemini 3.1 Pro (Preview)39.3
10Grok 4.2038.4
11Gemini 3 Flash34
12DeepSeek V4 Pro32.4
13MiniMax-M2.531.4
14Qwen 3 32B3.9
15Qwen 3 8B-43.3

Interactive version: theaggregate.ai/benchmark?slug=eibench-emotion-management-support · How It Works · Data refreshed daily, snapshot 2026-09-29.