CMMLU — leaderboard

CMMLU evaluates model capability on general knowledge tasks from the linked upstream source with Score as the primary reported metric.

Metric: 5-shot Avg Accuracy (%). Source: github.com. Status: saturated. 42 models tracked.

Top models

#ModelScore
1Qwen 2 72B89.65
2Qwen 1.5 110B88.32
3Qwen 2.5 72B85.67
4Qwen 1.5 72B83.54
5Qwen 1.5 32B82.25
6Qwen 1.5 7B72.5
7XuanYuan-70B71.1
8GPT-470.95
9Llama 3.1 70B Instruct64.38
10Qwen-7B58.66
11ChatGLM2 6B48.8
12Baichuan-7B44.43
13falcon-40B41.45
14LLaMA-65B39.8
15Llama 3 70B Instruct36.85

Interactive version: theaggregate.ai/benchmark?slug=cmmlu · How the rankings work · Data refreshed daily, snapshot 2026-07-22.