CommunityBench - Community Identification: leaderboard

Metric: Accuracy (%). Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1GPT-4o84.3
2DeepSeek V3 (0324)82.32
3Grok 482.23
4Llama 3.1 70B Instruct77.02
5Llama 3.3 70B Instruct73.8
6Qwen 2.5 72B Instruct68.96
7Qwen 2.5 14B Instruct68.53
8GLM-4 32B (0414)67.18
9Qwen 3 32B60.86
10Qwen 2.5 7B Instruct58.66
11GLM-4 9B (0414)51.25
12InternLM3-8B-Instruct49.24
13Mistral 7B Instruct (v0.3)48.55
14Qwen 3 14B48.52
15Llama 3.1 8B Instruct47.37

Interactive version: theaggregate.ai/benchmark?slug=communitybench-community-identification · How It Works · Data refreshed daily, snapshot 2026-09-25.