DialToM - Retrospective (Persuasion): leaderboard

Metric: Correctness rate (%) on four-option multiple-choice questions of the DialToM Retrospective task: identify a speaker's mental state (belief, desire, intention, emotion, knowledge, trust) from the dialogue context, persuasion domain (naturalistic human-human dialogues from PersuasionForGood; GPT-4o-drafted options verified by professional annotators and Dawid-Skene aggregation), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)98.78
2Qwen 3 235B A22B97.94
3GPT-595.28
4Gemini 2.5 Pro94.99
5Gemini 2.5 Flash94.4
6GPT-4.194.4
7Kimi K293.22
8Llama 3.3 70B87.91
9GPT-OSS-120B85.25
10Llama 4 Maverick79.28
11Mistral Small 3.274.34
12Mistral Nemo56.64
13Llama 3.1 8B34.51

Interactive version: theaggregate.ai/benchmark?slug=dialtom-retrospective-persuasion · How It Works · Data refreshed daily, snapshot 2026-10-07.