OpenEval - XSum — leaderboard

Metric: ROUGE-L. Source: huggingface.co. 19 models tracked.

Top models

#ModelScore
1Llama 2 70B33.64
2LLaMA-65B30.79
3Llama 2 13B30.43
4LLaMA-30B29.4
5LLaMA-13B27.47
6Llama 2 7B26.07
7LLaMA-7B23.11
8falcon-7B23.11
9falcon-40B Instruct22.84
10RedPajama-INCITE-Base-3B-v119.07
11RedPajama-INCITE-Instruct-3B-v118.46
12pythia-6.9B18.17
13falcon-7B Instruct13.41
14vicuna-13B-v1.39.27
15vicuna-7B-v1.30.84

Interactive version: theaggregate.ai/benchmark?slug=openeval-xsum · How the rankings work · Data refreshed daily, snapshot 2026-07-22.