EIBench (Emotion Management) - Defense: leaderboard
Metric: Anchor-based outcome score x 100 (-100 to +100): the simulated user's final emotion and relation state mapped per axis to 0 at the start anchor, +1 at the success anchor and -1 at the failure anchor (clipped), averaged over the 64 Defense scenarios (keeping a necessary boundary under user pressure while lowering frustration); 213 held-out test scenarios, multi-turn dialogue with an LLM user simulator, Qwen3-Max as the simulator; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 Max | -2.4 |
| 2 | Gemini 3.1 Pro (Preview) | -2.5 |
| 3 | DeepSeek V4 Pro | -7.8 |
| 4 | Kimi K2.6 | -8.8 |
| 5 | GPT-5.4 | -9.5 |
| 6 | Gemini 3 Flash | -11.1 |
| 7 | Grok 4.20 | -13.5 |
| 8 | Seed 2.0 Pro | -14.2 |
| 9 | Qwen 3 32B | -14.3 |
| 10 | Qwen 3.6 Max Preview | -14.3 |
| 11 | Claude Sonnet 4.6 | -14.6 |
| 12 | GLM-5.1 | -15.2 |
| 13 | MiniMax-M2.5 | -16.7 |
| 14 | Grok 4.20 (Reasoning) | -19 |
| 15 | Qwen 3 8B | -21.4 |
Interactive version: theaggregate.ai/benchmark?slug=eibench-emotion-management-defense · How It Works · Data refreshed daily, snapshot 2026-09-29.