SeaEval - Dialogue - DREAM (Zero-Shot) — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 24 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct96.28
2Qwen 2 72B Instruct96.13
3Llama 3.1 70B Instruct95.59
4Qwen 2.5 32B Instruct95.59
5Llama 3 70B Instruct94.81
6Qwen 2.5 14B Instruct94.61
7Gemma 2 9B (IT)94.17
8Qwen 2 7B Instruct93.53
9Qwen 2.5 7B Instruct93.48
10SeaLLMs-v3-7B-Chat92.65
11Llama 3.1 8B Instruct90.54
12Sailor2-8B-Chat90.54
13Qwen 2.5 3B Instruct90.3
14Llama 3 8B Instruct89.47
15Llama 3 8B Cpt Sea Lionv2.1 Instruct88.39

Interactive version: theaggregate.ai/benchmark?slug=seaeval-dialogue-dream-zero-shot · How the rankings work · Data refreshed daily, snapshot 2026-07-22.