SeaEval - Dialogue - SAMSum (Zero-Shot): leaderboard

Metric: Average ROUGE (0-100). Source: huggingface.co. 25 models tracked.

Top models

#ModelScore
1Gemma 2 2B (IT)31.12
2Gemma 2 9B (IT)31.01
3Llama 3 8B Cpt Sea Lionv2 Instruct30.7
4Llama 3 8B Cpt Sea Lionv2.1 Instruct30.5
5Qwen 2.5 7B Instruct29.88
6SeaLLMs-v3-7B-Chat29.6
7gemma2-9B-cpt-sea-lionv3-instruct29.51
8Llama 3 70B Instruct28.94
9Llama 3.1 70B Instruct28.93
10Qwen 2.5 72B Instruct28.85
11Llama 3 8B Instruct28.46
12Qwen 2.5 32B Instruct28.44
13Llama 3.1 8B Instruct28.21
14Llama 3.1 8B Cpt Sea Lionv3 Instruct28.07
15Qwen 2 72B Instruct28.01

Interactive version: theaggregate.ai/benchmark?slug=seaeval-dialogue-samsum-zero-shot · How It Works · Data refreshed daily, snapshot 2026-09-05.