CommunityBench - Community-Consistent Generation: leaderboard

Metric: BTL-Elo rating from majority-voted LLM-judge pairwise comparisons. Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.

Top models

#ModelScore
1DeepSeek R1 0528812.3
2Grok 4478.67
3Qwen 3 14B272.39
4GPT-4o256.28
5Qwen 3 32B237.8
6GLM-4 32B (0414)233.49
7DeepSeek V3 (0324)206.22
8GLM-4 9B (0414)176.85
9Qwen 3 8B87.77
10Llama 3.1 70B Instruct-54.9
11Llama 3.3 70B Instruct-56.64
12Qwen 2.5 72B Instruct-149.18
13Qwen 2.5 14B Instruct-154.87
14Llama 3.1 8B Instruct-166.28
15Qwen 2.5 7B Instruct-262.68

Interactive version: theaggregate.ai/benchmark?slug=communitybench-community-consistent-generation · How It Works · Data refreshed daily, snapshot 2026-09-25.