EuroEval Spanish Summarization - Mlsum ES — leaderboard

Metric: Score (%). Source: euroeval.com. 148 models tracked.

Top models

#ModelScore
1Bielik-11B-v2.3-Instruct31.19
2Gemma 3 27B (IT)29.68
3Qwen 2.5 72B Instruct29.62
4Apertus-70B-Instruct-250929.34
5granite-4.0-h-micro29.33
6Qwen 3 8B (Non-reasoning)29.13
7Qwen 3 14B (Non-reasoning)29.12
8Tiny Aya Global28.88
9Llama 3.1 8B Instruct28.87
10Llama 3.3 70B Instruct28.82
11Claude Sonnet 4.5 (Non-reasoning)28.69
12Qwen 3 30B A3B 2507 Instruct28.65
13GPT-4.128.52
14gemma-4-E2B-it28.48
15Gemini 2.5 Flash Lite (Non-reasoning)28.32

Interactive version: theaggregate.ai/benchmark?slug=euroeval-spanish-summarization-mlsum-es · How the rankings work · Data refreshed daily, snapshot 2026-07-22.