Persuasion (Lechmazur) — leaderboard

Multi-turn persuasion benchmark where one LLM tries to shift another model's stated position on policy propositions. 15 models, 15 topics, 6,296 conversations. Measures average signed stance shift toward persuader's assigned side.

Metric: Average Persuasion Strength. Source: github.com. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)1.71
2Claude Opus 4.6 (Thinking, High)1.67
3Seed 2.0 Pro1.64
4Claude Sonnet 4.6 (Thinking, High)1.58
5Kimi K2.5 (Thinking)1.33
6Gemini 3.1 Pro (Preview)1.23
7GLM-51.22
8Grok 4.20 Beta (0309) (Reasoning)1.2
9Qwen 3.5 397B A17B0.98
10MiniMax-M2.70.96
11Gemini 3.1 Flash Lite (Preview)0.81
12ERNIE 5.00.77
13DeepSeek V3.20.71
14MiMo-V2-Pro0.52
15Mistral Large 30.42

Interactive version: theaggregate.ai/benchmark?slug=persuasion-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.