DialToM - Prospective (Motivational Interviewing): leaderboard

Metric: Correctness rate (%) on four-option multiple-choice questions of the DialToM Prospective task: forecast the state-consistent next dialogue trajectory from an isolated mental-state profile with no dialogue context (the State-Driven Diagnostic Probe), motivational interviewing domain (naturalistic human-human dialogues from AnnoMI; GPT-4o-drafted options verified by professional annotators and Dawid-Skene aggregation), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)83.09
2Gemini 2.5 Pro38.24
3Gemini 2.5 Flash26.47
4Kimi K222.06
5Llama 3.3 70B18.42
6Qwen 3 235B A22B17.65
7GPT-514.71
8Llama 4 Maverick9.56
9GPT-OSS-120B7.35
10Llama 3.1 8B6.62
11Mistral Nemo6.62
12GPT-4.15.88
13Mistral Small 3.25.88

Interactive version: theaggregate.ai/benchmark?slug=dialtom-prospective-motivational-interviewing · How It Works · Data refreshed daily, snapshot 2026-10-07.