MOEVE - Summarization: leaderboard

Metric: Mean overall summarization score x 100 (0-100) over four German datasets (Eur-Lex-Sum, Swiss Leading Decision Summarization, KIKC Summary, German Ministry Publications), combining BERTScore, SemScore, LLM-judged factual correctness and the German-output proportion; German prompts, default decoding, open-weight models on vLLM or Ollama at the quantization listed in the paper's model table, others by API; the paper prints only the top 10 of 39 evaluated models; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Mistral Large 378
2GPT-4o Mini77
3Phi-477
4Mistral Small 3.176.7
5Llama 3.3 70B76.6
6GPT-4o76.3
7DeepSeek R176.1
8GPT-OSS-120B75.8

Interactive version: theaggregate.ai/benchmark?slug=moeve-summarization · How It Works · Data refreshed daily, snapshot 2026-09-29.