OpenCompass LLM - Reasoning — leaderboard

Metric: Score (%). Source: rank.opencompass.org.cn. 23 models tracked.

Top models

#ModelScore
1GPT-5.4 (High)64.4
2Claude Opus 4.7 (High)63.5
3Kimi K2.662.2
4Claude Sonnet 4.6 (High)60.6
5DeepSeek V4 Pro60
6Gemini 3.1 Pro (Preview)59.7
7Qwen 3.6 Max Preview59.2
8Hy3-preview (High)58.5
9DeepSeek V4 Flash57.2
10GPT-5.4 Mini (High)55.9
11Step 3.5 Flash55.6
12Gemma 4 31B (IT)54.6
13Qwen 3.5 397B A17B54.5
14GLM-5.154.3
15Qwen 3.6 27B52.7

Interactive version: theaggregate.ai/benchmark?slug=opencompass-llm-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.