OpenCompass Language - Chinese Semantic Understanding - Chinese (CompassBench 2404): leaderboard

Metric: Score (%). Source: rank.opencompass.org.cn. 31 models tracked.

Top models

#ModelScore
1OrionStar-Yi-34B-Chat71
2Qwen-72B-Chat63.7
3Qwen-14B-Chat63.3
4GPT-4 Turbo (Preview)59
5Yi 34B (Chat)56.3
6Qwen-7B Chat49
7Baichuan2-13B-Chat44
8Yi 6B (Chat)41
9Baichuan2-7B-Chat26.7
10Mistral 7B Instruct (v0.2)25
11WizardLM-70B-V1.024.7
12deepseek-llm-7B-chat22.7
13WizardLM-13B-V1.214
14Llama 2 70B Chat13.3
15Llama 2 7B Chat11

Interactive version: theaggregate.ai/benchmark?slug=opencompass-language-chinese-semantic-understanding-chinese-compassbench-2404 · How It Works · Data refreshed daily, snapshot 2026-09-19.