SeaEval - Multilingual Reasoning - CMMLU (Zero-Shot) — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 22 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct83.44
2Qwen 2 72B Instruct82.94
3Qwen 2.5 32B Instruct82.73
4Qwen 2.5 14B Instruct78.08
5Qwen 2 7B Instruct77.28
6SeaLLMs-v3-7B-Chat76.84
7Qwen 2.5 7B Instruct74.87
8Llama 3.1 70B Instruct68.15
9Qwen 2.5 3B Instruct66.21
10Llama 3 70B Instruct64.95
11Sailor2-8B-Chat64.17
12Qwen 2.5 1.5B Instruct59.76
13Gemma 2 9B (IT)57
14Llama 3.1 8B Instruct51.83
15Llama 3 8B Cpt Sea Lionv2.1 Instruct49.22

Interactive version: theaggregate.ai/benchmark?slug=seaeval-multilingual-reasoning-cmmlu-zero-shot · How the rankings work · Data refreshed daily, snapshot 2026-07-22.