SeaEval - Multilingual Reasoning - C-Eval (Zero-Shot): leaderboard

Metric: Accuracy (%). Source: huggingface.co. 24 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct83.25
2Qwen 2 72B Instruct83.13
3Qwen 2.5 32B Instruct82.63
4Qwen 2.5 14B Instruct78.39
5SeaLLMs-v3-7B-Chat76.59
6Qwen 2 7B Instruct76.15
7Qwen 2.5 7B Instruct74.6
8Llama 3.1 70B Cpt Sea Lionv3 Instruct67.62
9Llama 3.1 70B Instruct66.13
10Qwen 2.5 3B Instruct65.38
11Llama 3 70B Instruct62.2
12Qwen 2.5 1.5B Instruct59.71
13Sailor2-8B-Chat59.46
14gemma2-9B-cpt-sea-lionv3-instruct57.22
15Gemma 2 9B (IT)55.23

Interactive version: theaggregate.ai/benchmark?slug=seaeval-multilingual-reasoning-c-eval-zero-shot · How It Works · Data refreshed daily, snapshot 2026-09-05.