EIBench (Emotion Management) - Repair: leaderboard

Metric: Anchor-based outcome score x 100 (-100 to +100): the simulated user's final emotion and relation state mapped per axis to 0 at the start anchor, +1 at the success anchor and -1 at the failure anchor (clipped), averaged over the 52 Repair scenarios (acknowledging a mistake and restoring trust); 213 held-out test scenarios, multi-turn dialogue with an LLM user simulator, Qwen3-Max as the simulator; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 15 models tracked.

Top models

#ModelScore
1GLM-5.119.3
2Gemini 3.1 Pro (Preview)18.6
3Gemini 3 Flash18.1
4Claude Sonnet 4.617.9
5Kimi K2.617.6
6Qwen 3.6 Max Preview14.7
7GPT-5.414.5
8Qwen 3 Max14.1
9DeepSeek V4 Pro12.2
10Grok 4.20 (Reasoning)12.2
11Seed 2.0 Pro10.7
12Qwen 3 32B6.8
13MiniMax-M2.55.2
14Grok 4.205.1
15Qwen 3 8B-35.9

Interactive version: theaggregate.ai/benchmark?slug=eibench-emotion-management-repair · How It Works · Data refreshed daily, snapshot 2026-09-29.