EIBench (Emotion Management) - Charm: leaderboard

Metric: Anchor-based outcome score x 100 (-100 to +100): the simulated user's final emotion and relation state mapped per axis to 0 at the start anchor, +1 at the success anchor and -1 at the failure anchor (clipped), averaged over the 46 Charm scenarios (building rapport with a stranger; the model speaks first); 213 held-out test scenarios, multi-turn dialogue with an LLM user simulator, Qwen3-Max as the simulator; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.643.8
2GPT-5.441.5
3Qwen 3.6 Max Preview39.1
4GLM-5.137.9
5Kimi K2.637.3
6Gemini 3 Flash37
7Qwen 3 Max32
8Seed 2.0 Pro31.9
9Gemini 3.1 Pro (Preview)31.3
10Grok 4.2031
11DeepSeek V4 Pro28
12Grok 4.20 (Reasoning)28
13Qwen 3 32B23.7
14MiniMax-M2.521.6
15Qwen 3 8B14.6

Interactive version: theaggregate.ai/benchmark?slug=eibench-emotion-management-charm · How It Works · Data refreshed daily, snapshot 2026-09-29.