OpenCompass Reasoning - Inductive - English (CompassBench 2404): leaderboard

Metric: Score (%). Source: rank.opencompass.org.cn. 31 models tracked.

Top models

#ModelScore
1GPT-4 Turbo (Preview)47.9
2OrionStar-Yi-34B-Chat34.2
3Baichuan2-13B-Chat32.3
4Yi 34B (Chat)31.7
5deepseek-llm-7B-chat29.2
6Qwen-7B Chat23.2
7Qwen-72B-Chat21.8
8Baichuan2-7B-Chat21.8
9Qwen-14B-Chat20.8
10Llama 2 70B Chat20
11Yi 6B (Chat)17.7
12Llama 2 13B Chat17.5
13Llama 2 7B Chat16.3
14WizardLM-13B-V1.215.7
15WizardLM-70B-V1.08.2

Interactive version: theaggregate.ai/benchmark?slug=opencompass-reasoning-inductive-english-compassbench-2404 · How It Works · Data refreshed daily, snapshot 2026-09-19.