SeaEval - Multilingual Reasoning - C-Eval (Zero-Shot) — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 24 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct83.25
2Qwen 2 72B Instruct83.13
3Qwen 2.5 32B Instruct82.63
4Qwen 2.5 14B Instruct78.39
5SeaLLMs-v3-7B-Chat76.59
6Qwen 2 7B Instruct76.15
7Qwen 2.5 7B Instruct74.6
8Llama 3.1 70B Instruct66.13
9Qwen 2.5 3B Instruct65.38
10Llama 3 70B Instruct62.2
11Qwen 2.5 1.5B Instruct59.71
12Sailor2-8B-Chat59.46
13Gemma 2 9B (IT)55.23
14Llama 3.1 8B Instruct51.81
15Llama 3 8B Cpt Sea Lionv2.1 Instruct50.19

Interactive version: theaggregate.ai/benchmark?slug=seaeval-multilingual-reasoning-c-eval-zero-shot · How the rankings work · Data refreshed daily, snapshot 2026-07-22.