SeaEval - Dialogue - DialogSum (Zero-Shot) — leaderboard

Metric: Average ROUGE (0-100). Source: huggingface.co. 24 models tracked.

Top models

#ModelScore
1Gemma 2 2B (IT)25.97
2Gemma 2 9B (IT)25.61
3Llama 3 70B Instruct25.57
4Llama 3 8B Cpt Sea Lionv2.1 Instruct25.38
5Llama 3.1 70B Instruct25.26
6Llama 3.1 8B Instruct25.25
7Qwen 2.5 7B Instruct25.03
8SeaLLMs-v3-7B-Chat24.89
9Llama 3 8B Instruct23.98
10Qwen 2.5 32B Instruct23.94
11Qwen 2.5 72B Instruct23.46
12Qwen 2.5 14B Instruct23.43
13Qwen 2.5 3B Instruct22.11
14Qwen 2 72B Instruct21.83
15Qwen 2 7B Instruct20.93

Interactive version: theaggregate.ai/benchmark?slug=seaeval-dialogue-dialogsum-zero-shot · How the rankings work · Data refreshed daily, snapshot 2026-07-22.