Multi-turn Debate (Lechmazur): leaderboard

Adversarial multi-turn debate benchmark where LLMs argue opposing sides of policy propositions. Each matchup runs twice with sides reversed. Ranked via Bradley-Terry ratings across 1,000+ debates judged by LLMs. Tests factual accuracy under pressure, counterargument quality, and epistemic discipline.

Metric: Bradley-Terry Rating. Source: github.com. Status: years away from saturation. 46 models tracked.

Top models

#ModelScore
1Claude Opus 5 (High)1748.5
2Claude Fable 5 (High)1747
3Kimi K31723.1
4Claude Opus 4.7 (High)1672.6
5Muse Spark 1.1 (High)1668
6GPT-5.6 Sol (High)1664
7Claude Opus 4.8 (High)1655.1
8Grok 4.6 (High)1628.4
9Claude Sonnet 5 (High)1610.8
10GLM-5.2 (Max)1585.2
11Claude Sonnet 4.6 (High)1585.2
12Qwen 3.8 Max1579.4
13GPT-5.4 (High)1571.7
14GPT-5.5 (High)1551.7
15GLM-5.11542.8

Interactive version: theaggregate.ai/benchmark?slug=multi-turn-debate-lechmazur · How It Works · Data refreshed daily, snapshot 2026-09-05.