Persuasion (Lechmazur) — leaderboard
Multi-turn persuasion benchmark where one LLM tries to shift another model's stated position on policy propositions. 15 models, 15 topics, 6,296 conversations. Measures average signed stance shift toward persuader's assigned side.
Metric: Average Persuasion Strength. Source: github.com. Status: saturation imminent. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 (High) | 1.71 |
| 2 | Claude Opus 4.6 (Thinking, High) | 1.67 |
| 3 | Seed 2.0 Pro | 1.64 |
| 4 | Claude Sonnet 4.6 (Thinking, High) | 1.58 |
| 5 | Kimi K2.5 (Thinking) | 1.33 |
| 6 | Gemini 3.1 Pro (Preview) | 1.23 |
| 7 | GLM-5 | 1.22 |
| 8 | Grok 4.20 Beta (0309) (Reasoning) | 1.2 |
| 9 | Qwen 3.5 397B A17B | 0.98 |
| 10 | MiniMax-M2.7 | 0.96 |
| 11 | Gemini 3.1 Flash Lite (Preview) | 0.81 |
| 12 | ERNIE 5.0 | 0.77 |
| 13 | DeepSeek V3.2 | 0.71 |
| 14 | MiMo-V2-Pro | 0.52 |
| 15 | Mistral Large 3 | 0.42 |
Interactive version: theaggregate.ai/benchmark?slug=persuasion-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.