EIBench (Emotion Management) - Human Arena: leaderboard

Metric: Elo rating from more than 4,000 human pairwise judgments: annotators converse with two anonymized models under the same EIBench scenario profile and pick the one that handles it better; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)1616
2Gemini 3 Flash1606
3GPT-5.41593
4Claude Sonnet 4.61548
5Seed 2.0 Pro1536
6Qwen 3.6 Max Preview1519
7GLM-5.11504
8Qwen 3 Max1487
9DeepSeek V4 Pro1484
10Kimi K2.61478
11Grok 4.201476
12MiniMax-M2.51471
13Grok 4.20 (Reasoning)1458
14Qwen 3 32B1405
15Qwen 3 8B1316

Interactive version: theaggregate.ai/benchmark?slug=eibench-emotion-management-human-arena · How It Works · Data refreshed daily, snapshot 2026-09-29.