Open Chinese LLM Leaderboard — leaderboard

Chinese language evaluation across C-ARC, C-HellaSwag, C-TruthfulQA, C-Winogrande, C-GSM8K, CMMLU, and Chinese semantic understanding.

Metric: Average Score (%). Source: huggingface.co. Status: saturation imminent. 339 models tracked.

Top models

#ModelScore
1Qwen 2 72B Instruct74.88
2TW3-JRGL-v274.66
3free-evo-qwen72B-v0.8-re74.62
4Rhea-72B-v0.574.34
5Le_Triomphant-ECE-TW373.83
6MultiVerse_70B73.8
7Smaug-72B-v0.173.16
8Qwen 2 72B72.66
9luxia-21.4B-alignment-v1.071.65
10ECE-TW3-JRGL-V571.22
11YiSM-blossom5.1-34B-SLERP70.1
12Higgs-Llama-3-70B70
13YiSM-34B-0rn69.66
14Yi-34Bx2-MoE-60B-DPO69.46
15luxia-21.4B-alignment-v1.269.42

Interactive version: theaggregate.ai/benchmark?slug=open-chinese-llm-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.