SuperCLUE — leaderboard

Comprehensive Chinese LLM evaluation covering basic capabilities, Chinese-specific features, and academic proficiency. 20+ models ranked across 10 sub-tasks.

Metric: Total Score. Source: www.cluebenchmarks.com. Status: saturated. 20 models tracked.

Top models

#ModelScore
1GPT-470.89
2Claude 260.94
3GPT-3.5 Turbo59.79
4ChatGLM2 6B53.85
5Llama 2 13B Chat34.26
6vicuna-13B31.7

Interactive version: theaggregate.ai/benchmark?slug=superclue · How the rankings work · Data refreshed daily, snapshot 2026-07-22.