SeaEval - Cultural Reasoning - PH-Eval (Zero-Shot) — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 25 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct72
2Qwen 2.5 32B Instruct70
3Llama 3.1 70B Instruct68
4Llama 3 70B Instruct63
5Qwen 2 72B Instruct62
6Qwen 2.5 14B Instruct60
7Gemma 2 9B (IT)58
8Llama 3 8B Instruct58
9Llama 3.1 8B Instruct57
10Llama 3 8B Cpt Sea Lionv2.1 Instruct56
11Qwen 2.5 7B Instruct55
12Sailor2-8B-Chat53
13Qwen 2 7B Instruct52
14SeaLLMs-v3-7B-Chat47
15Gemma 2 2B (IT)40

Interactive version: theaggregate.ai/benchmark?slug=seaeval-cultural-reasoning-ph-eval-zero-shot · How the rankings work · Data refreshed daily, snapshot 2026-07-22.