EIBench (Emotion Management): leaderboard

Metric: Anchor-based outcome score x 100 (-100 to +100): the simulated user's final emotion and relation state mapped per axis to 0 at the start anchor, +1 at the success anchor and -1 at the failure anchor (clipped), scenario-weighted average over the four scene types; 213 held-out test scenarios, multi-turn dialogue with an LLM user simulator, Qwen3-Max as the simulator; higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.623
2GPT-5.421.4
3Qwen 3 Max20.6
4Gemini 3.1 Pro (Preview)20
5Kimi K2.619.8
6GLM-5.119.1
7Qwen 3.6 Max Preview17.9
8Gemini 3 Flash17.2
9Seed 2.0 Pro14.8
10DeepSeek V4 Pro14.4
11Grok 4.2013.1
12Grok 4.20 (Reasoning)13.1
13MiniMax-M2.58.4
14Qwen 3 32B3.4
15Qwen 3 8B-22.4

Interactive version: theaggregate.ai/benchmark?slug=eibench-emotion-management · How It Works · Data refreshed daily, snapshot 2026-09-29.