OpenCompass Language - Intention Recognition - English (CompassBench 2404): leaderboard

Metric: Score (%). Source: rank.opencompass.org.cn. 31 models tracked.

Top models

#ModelScore
1GPT-4 Turbo (Preview)77
2Yi 34B (Chat)70.3
3Qwen-72B-Chat70.3
4OrionStar-Yi-34B-Chat67
5Qwen-14B-Chat66
6Llama 2 70B Chat61
7Qwen-7B Chat57
8Yi 6B (Chat)51.7
9Baichuan2-13B-Chat50
10Mistral 7B Instruct (v0.2)46.3
11Llama 2 13B Chat46
12Baichuan2-7B-Chat41.7
13Llama 2 7B Chat32.3
14deepseek-llm-7B-chat28
15WizardLM-13B-V1.226.3

Interactive version: theaggregate.ai/benchmark?slug=opencompass-language-intention-recognition-english-compassbench-2404 · How It Works · Data refreshed daily, snapshot 2026-09-19.