EuroEval Polish Summarization - PSC — leaderboard

Metric: Score (%). Source: euroeval.com. 148 models tracked.

Top models

#ModelScore
1Ministral-3-3B-Instruct-251225.82
2Ministral-3-8B-Instruct-251225.67
3GPT-5.4 Nano (Non-reasoning)25.44
4GPT-OSS-20B (Low)25.36
5GPT-OSS-20B (Medium)25.35
6GPT-5.4 Nano (High)25.24
7GPT-5.4 Nano (Medium)25.21
8LFM2 1.2B25.18
9Apertus-70B-Instruct-250925.01
10Ministral-3-14B-Instruct-251224.73
11Bielik-11B-v2.3-Instruct24.62
12Claude Haiku 4.5 (20251001)24.61
13Qwen 3 1.7B (Non-reasoning)24.52
14GPT-5 Nano24.36
15Qwen 3 30B A3B 2507 Instruct24.19

Interactive version: theaggregate.ai/benchmark?slug=euroeval-polish-summarization-psc · How the rankings work · Data refreshed daily, snapshot 2026-07-22.