Chinese Competitive Debating - Best Debater: leaderboard

Metric: Best-debater accuracy (%; the model distributes nine best-debater votes among the match's debaters from the full transcript; correct when its top debater is among those with the most judge votes; 148 competitive Chinese-language debate matches (4v4 and 2v2 formats), each adjudicated by three professional judges under a shared rubric; zero-shot on manually verified transcripts with fixed prompts; random choice 18.3). Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (Medium)56.8
2DeepSeek V4 Pro55.4
3Claude Opus 5 (Medium)52
4DeepSeek V4 Flash (Reasoning)43.9
5DeepSeek V4 Flash43.2
6GPT-5.6 Luna (Medium)43.2
7Claude Sonnet 5 (Medium)41.9
8Gemini 3.5 Flash Lite (Medium)34.5

Interactive version: theaggregate.ai/benchmark?slug=chinese-competitive-debating-best-debater · How It Works · Data refreshed daily, snapshot 2026-09-26.