DialToM - Prospective (Emotional Support): leaderboard

Metric: Correctness rate (%) on four-option multiple-choice questions of the DialToM Prospective task: forecast the state-consistent next dialogue trajectory from an isolated mental-state profile with no dialogue context (the State-Driven Diagnostic Probe), emotional support domain (naturalistic human-human dialogues from ESConv; GPT-4o-drafted options verified by professional annotators and Dawid-Skene aggregation), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)80.46
2Gemini 2.5 Pro29.06
3Qwen 3 235B A22B29.06
4Kimi K225.62
5Gemini 2.5 Flash25.12
6Llama 3.3 70B19.7
7Llama 3.1 8B16.75
8GPT-516.26
9GPT-OSS-120B15.76
10Mistral Nemo14.29
11GPT-4.112.81
12Llama 4 Maverick11.82
13Mistral Small 3.24.93

Interactive version: theaggregate.ai/benchmark?slug=dialtom-prospective-emotional-support · How It Works · Data refreshed daily, snapshot 2026-10-07.