OpenCompass Language - Dialogue - Chinese (CompassBench 2409): leaderboard

Metric: Score (%). Source: rank.opencompass.org.cn. 30 models tracked.

Top models

#ModelScore
1GLM-4 Plus67
2DeepSeek V2.563.5
3Step 2 16K61.7
4GPT-4o (2024-05-13)57.5
5Yi-1.5-9B Chat52.4
6Mistral Large 2 (Jul)52
7Qwen 2.5 72B Instruct50.4
8GPT-4o Mini (2024-07-18)48.9
9GLM-4 9B Chat46
10Gemini 1.5 Pro45.1
11GPT-4o (2024-08-06)44.9
12Yi Large44.9
13Mistral-Small-Instruct-240943.8
14Claude 3.5 Sonnet (20240620)40.7
15Mistral Nemo Instruct (2407)39.8

Interactive version: theaggregate.ai/benchmark?slug=opencompass-language-dialogue-chinese-compassbench-2409 · How It Works · Data refreshed daily, snapshot 2026-09-19.