SeaEval - Multilingual Reasoning - IndoMMLU (Zero-Shot) — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 21 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct67.41
2Qwen 2 72B Instruct63.86
3Qwen 2.5 72B Instruct63.81
4Llama 3 70B Instruct63.24
5Qwen 2.5 32B Instruct63.15
6Gemma 2 9B (IT)60.7
7Qwen 2.5 14B Instruct60.1
8Llama 3.1 8B Instruct56.05
9Qwen 2.5 7B Instruct56.01
10Qwen 2 7B Instruct53.86
11Llama 3 8B Cpt Sea Lionv2.1 Instruct52.69
12SeaLLMs-v3-7B-Chat52.67
13Llama 3 8B Instruct52.65
14Qwen 2.5 3B Instruct49.66
15Gemma 2 2B (IT)48.22

Interactive version: theaggregate.ai/benchmark?slug=seaeval-multilingual-reasoning-indommlu-zero-shot · How the rankings work · Data refreshed daily, snapshot 2026-07-22.