BlueBench - Summarization — leaderboard

Metric: Score (%). Source: huggingface.co. 18 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct19.47
2GPT-4.118.25
3Mistral Large18.24
4Llama 3.2 3B Instruct17.99
5GPT-4.1 Mini17.92
6Llama 3.2 1B Instruct17.62
7Mistral Medium 317.43
8Pixtral-12B17.28
9GPT-4o17.06
10GPT-4.1 Nano16.94
11O116.77
12O3 Mini16.66
13O4 Mini16.3

Interactive version: theaggregate.ai/benchmark?slug=bluebench-summarization · How the rankings work · Data refreshed daily, snapshot 2026-07-22.