CLUE — leaderboard
Chinese Language Understanding Evaluation benchmark covering Chinese classification and reading-comprehension tasks with public leaderboard baselines.
Metric: Classification Score (%). Source: github.com. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | RoBERTa-wwm-large | 72.83 |
| 2 | RoBERTa-large | 71.7 |
| 3 | ALBERT-xxlarge | 71.04 |
| 4 | RoBERTa-wwm-ext | 70.1 |
| 5 | ALBERT-xlarge | 68.92 |
| 6 | BERT-base | 68.77 |
| 7 | BERT-wwm-ext | 68.75 |
| 8 | XLNet-mid | 68.58 |
| 9 | ERNIE-base | 68.55 |
| 10 | ALBERT-large | 67.91 |
Interactive version: theaggregate.ai/benchmark?slug=clue · How the rankings work · Data refreshed daily, snapshot 2026-07-22.