Adversarial Empathy Benchmark (NoThink Mode): leaderboard

Metric: Final score (0-1; the simulated user's emotion after eight turns divided by 100, mean over 60 adversarial dialogues in six trajectory types; Mistral-7B-Instruct-v0.3 SAGE simulator; standard mode, no reasoning scaffold). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1Qwen 2.5 7B Instruct0.8
2Qwen 2.5 1.5B Instruct0.75

Interactive version: theaggregate.ai/benchmark?slug=adversarial-empathy-benchmark-nothink-mode · How It Works · Data refreshed daily, snapshot 2026-09-26.