DialToM - Retrospective (Emotional Support): leaderboard

Metric: Correctness rate (%) on four-option multiple-choice questions of the DialToM Retrospective task: identify a speaker's mental state (belief, desire, intention, emotion, knowledge, trust) from the dialogue context, emotional support domain (naturalistic human-human dialogues from ESConv; GPT-4o-drafted options verified by professional annotators and Dawid-Skene aggregation), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)99.15
2Gemini 2.5 Pro97.18
3Qwen 3 235B A22B97.18
4Gemini 2.5 Flash95.48
5Kimi K295.48
6GPT-595.2
7GPT-4.194.07
8GPT-OSS-120B88.98
9Llama 4 Maverick87.57
10Llama 3.3 70B84.75
11Mistral Small 3.268.64
12Mistral Nemo52.26
13Llama 3.1 8B27.12

Interactive version: theaggregate.ai/benchmark?slug=dialtom-retrospective-emotional-support · How It Works · Data refreshed daily, snapshot 2026-10-07.