Open Chinese LLM Leaderboard — leaderboard
Chinese language evaluation across C-ARC, C-HellaSwag, C-TruthfulQA, C-Winogrande, C-GSM8K, CMMLU, and Chinese semantic understanding.
Metric: Average Score (%). Source: huggingface.co. Status: saturation imminent. 339 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2 72B Instruct | 74.88 |
| 2 | TW3-JRGL-v2 | 74.66 |
| 3 | free-evo-qwen72B-v0.8-re | 74.62 |
| 4 | Rhea-72B-v0.5 | 74.34 |
| 5 | Le_Triomphant-ECE-TW3 | 73.83 |
| 6 | MultiVerse_70B | 73.8 |
| 7 | Smaug-72B-v0.1 | 73.16 |
| 8 | Qwen 2 72B | 72.66 |
| 9 | luxia-21.4B-alignment-v1.0 | 71.65 |
| 10 | ECE-TW3-JRGL-V5 | 71.22 |
| 11 | YiSM-blossom5.1-34B-SLERP | 70.1 |
| 12 | Higgs-Llama-3-70B | 70 |
| 13 | YiSM-34B-0rn | 69.66 |
| 14 | Yi-34Bx2-MoE-60B-DPO | 69.46 |
| 15 | luxia-21.4B-alignment-v1.2 | 69.42 |
Interactive version: theaggregate.ai/benchmark?slug=open-chinese-llm-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.