Psi-Bench: leaderboard

Metric: Average persuasion effect over the three scenarios (1-9; DeepSeek-v3.2 judge, which also plays the persona-profiled client; three rounds after the client's opening message, client profile hidden from the model). Source: arxiv.org. Saturation forecast: Around 2029. 10 models tracked.

Top models

#ModelScore
1GPT-5.15.79
2Qwen 3 Next 80B A3B5.38
3DeepSeek V4 Pro5.22
4Gemini 3.1 Pro (Preview)5.11
5GPT-5 Mini4.92
6Gemini 3 Flash4.89
7Qwen 3 32B4.54
8DeepSeek V3.24.47
9Grok 4 Fast4.19
10Qwen 3 8B4.05

Interactive version: theaggregate.ai/benchmark?slug=psi-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.