Chinese Competitive Debating - Winner Prediction: leaderboard
Metric: Accuracy (%; the model reads a full match transcript and gives the probability that the affirmative side wins; correct when it picks the side the judges' majority chose, an output of exactly 0.5 counts as wrong; 148 competitive Chinese-language debate matches (4v4 and 2v2 formats), each adjudicated by three professional judges under a shared rubric; zero-shot on manually verified transcripts with fixed prompts; always-affirmative baseline 50.7). Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 (Medium) | 66.2 |
| 2 | DeepSeek V4 Pro | 61.5 |
| 3 | GPT-5.6 Sol (Medium) | 60.1 |
| 4 | DeepSeek V4 Flash (Reasoning) | 60.1 |
| 5 | Claude Sonnet 5 (Medium) | 59.5 |
| 6 | Gemini 3.5 Flash Lite (Medium) | 59.5 |
| 7 | DeepSeek V4 Flash | 58.8 |
| 8 | GPT-5.6 Luna (Medium) | 56.1 |
Interactive version: theaggregate.ai/benchmark?slug=chinese-competitive-debating-winner-prediction · How It Works · Data refreshed daily, snapshot 2026-09-26.