EIBench (Emotion Management) - Defense: leaderboard

Metric: Anchor-based outcome score x 100 (-100 to +100): the simulated user's final emotion and relation state mapped per axis to 0 at the start anchor, +1 at the success anchor and -1 at the failure anchor (clipped), averaged over the 64 Defense scenarios (keeping a necessary boundary under user pressure while lowering frustration); 213 held-out test scenarios, multi-turn dialogue with an LLM user simulator, Qwen3-Max as the simulator; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 15 models tracked.

Top models

#ModelScore
1Qwen 3 Max-2.4
2Gemini 3.1 Pro (Preview)-2.5
3DeepSeek V4 Pro-7.8
4Kimi K2.6-8.8
5GPT-5.4-9.5
6Gemini 3 Flash-11.1
7Grok 4.20-13.5
8Seed 2.0 Pro-14.2
9Qwen 3 32B-14.3
10Qwen 3.6 Max Preview-14.3
11Claude Sonnet 4.6-14.6
12GLM-5.1-15.2
13MiniMax-M2.5-16.7
14Grok 4.20 (Reasoning)-19
15Qwen 3 8B-21.4

Interactive version: theaggregate.ai/benchmark?slug=eibench-emotion-management-defense · How It Works · Data refreshed daily, snapshot 2026-09-29.