Deception Effectiveness (Lechmazur) — leaderboard
Measures LLM capability to generate persuasive disinformation: models craft misleading arguments for incorrect answers to fact-based questions.
Metric: Deception Score. Source: github.com. Status: saturation imminent. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.5 Sonnet | 1.1 |
| 2 | Mistral Large 2 (Jul) | 1.09 |
| 3 | O1 Preview | 1.03 |
| 4 | Grok 2 | 0.96 |
| 5 | Gemini 1.5 Pro (Sept) | 0.94 |
| 6 | Llama 3.1 405B | 0.78 |
| 7 | Llama 3.1 70B | 0.71 |
| 8 | O1 Mini | 0.67 |
| 9 | Claude 3 Haiku | 0.66 |
| 10 | Claude 3 Opus | 0.65 |
| 11 | DeepSeek V2.5 | 0.61 |
| 12 | Gemini 1.5 Flash | 0.61 |
| 13 | GPT-4o | 0.6 |
| 14 | GPT-4 Turbo | 0.56 |
| 15 | Gemma 2 27B | 0.52 |
Interactive version: theaggregate.ai/benchmark?slug=deception-effectiveness-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.