DialToM - Prospective (Persuasion): leaderboard

Metric: Correctness rate (%) on four-option multiple-choice questions of the DialToM Prospective task: forecast the state-consistent next dialogue trajectory from an isolated mental-state profile with no dialogue context (the State-Driven Diagnostic Probe), persuasion domain (naturalistic human-human dialogues from PersuasionForGood; GPT-4o-drafted options verified by professional annotators and Dawid-Skene aggregation), zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)86.96
2Gemini 2.5 Pro19.57
3Gemini 2.5 Flash17.93
4Qwen 3 235B A22B17.93
5GPT-516.85
6Kimi K213.04
7GPT-4.110.87
8Llama 3.1 8B10.33
9Mistral Small 3.27.6
10Llama 3.3 70B7.07
11GPT-OSS-120B6.52
12Llama 4 Maverick5.6
13Mistral Nemo5.44

Interactive version: theaggregate.ai/benchmark?slug=dialtom-prospective-persuasion · How It Works · Data refreshed daily, snapshot 2026-10-07.