OpenCompass Reasoning - Common — leaderboard

Metric: Score (%). Source: rank.opencompass.org.cn. 23 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)76.8
2DeepSeek V4 Pro76
3Claude Opus 4.7 (High)76
4Gemini 3.1 Pro (Preview)75.8
5Kimi K2.675.5
6DeepSeek V4 Flash75.5
7Claude Sonnet 4.6 (High)75.5
8Qwen 3.6 Max Preview75
9Hy3-preview (High)73.5
10Qwen 3.5 397B A17B72.5
11Step 3.5 Flash72.5
12Gemma 4 31B (IT)71.2
13GPT-5.4 Mini (High)70.2
14GLM-5.169.9
15Gemini 3.1 Flash Lite (Preview)69.6

Interactive version: theaggregate.ai/benchmark?slug=opencompass-reasoning-common · How the rankings work · Data refreshed daily, snapshot 2026-07-22.