SuperCLUE — leaderboard
Comprehensive Chinese LLM evaluation covering basic capabilities, Chinese-specific features, and academic proficiency. 20+ models ranked across 10 sub-tasks.
Metric: Total Score. Source: www.cluebenchmarks.com. Status: saturated. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 | 70.89 |
| 2 | Claude 2 | 60.94 |
| 3 | GPT-3.5 Turbo | 59.79 |
| 4 | ChatGLM2 6B | 53.85 |
| 5 | Llama 2 13B Chat | 34.26 |
| 6 | vicuna-13B | 31.7 |
Interactive version: theaggregate.ai/benchmark?slug=superclue · How the rankings work · Data refreshed daily, snapshot 2026-07-22.