OpenEval - XSum — leaderboard
Metric: ROUGE-L. Source: huggingface.co. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 2 70B | 33.64 |
| 2 | LLaMA-65B | 30.79 |
| 3 | Llama 2 13B | 30.43 |
| 4 | LLaMA-30B | 29.4 |
| 5 | LLaMA-13B | 27.47 |
| 6 | Llama 2 7B | 26.07 |
| 7 | LLaMA-7B | 23.11 |
| 8 | falcon-7B | 23.11 |
| 9 | falcon-40B Instruct | 22.84 |
| 10 | RedPajama-INCITE-Base-3B-v1 | 19.07 |
| 11 | RedPajama-INCITE-Instruct-3B-v1 | 18.46 |
| 12 | pythia-6.9B | 18.17 |
| 13 | falcon-7B Instruct | 13.41 |
| 14 | vicuna-13B-v1.3 | 9.27 |
| 15 | vicuna-7B-v1.3 | 0.84 |
Interactive version: theaggregate.ai/benchmark?slug=openeval-xsum · How the rankings work · Data refreshed daily, snapshot 2026-07-22.