DialToM - Retrospective (Motivational Interviewing): leaderboard

Metric: Correctness rate (%) on four-option multiple-choice questions of the DialToM Retrospective task: identify a speaker's mental state (belief, desire, intention, emotion, knowledge, trust) from the dialogue context, motivational interviewing domain (naturalistic human-human dialogues from AnnoMI; GPT-4o-drafted options verified by professional annotators and Dawid-Skene aggregation), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)97.71
2Qwen 3 235B A22B95.1
3Gemini 2.5 Flash94.77
4GPT-593.46
5GPT-4.193.14
6Gemini 2.5 Pro92.48
7Kimi K292.16
8Llama 4 Maverick87.58
9Llama 3.3 70B77.45
10GPT-OSS-120B75.49
11Mistral Small 3.261.11
12Mistral Nemo42.16
13Llama 3.1 8B30.39

Interactive version: theaggregate.ai/benchmark?slug=dialtom-retrospective-motivational-interviewing · How It Works · Data refreshed daily, snapshot 2026-10-07.