Open PL LLM Leaderboard — leaderboard

Polish language LLM evaluation across Polemo2, 8tags, Belebele, Polish QA, EQ-Bench, RAG, and other Polish understanding and generation tasks.

Metric: Average Score (%). Source: huggingface.co. Status: saturated. 287 models tracked.

Top models

#ModelScore
1Mistral Large 2 (Nov) Instruct (2411)69.84
2Llama 3.1 405B Instruct FP869.44
3Mistral Large 2 (Jul)69.11
4Qwen 2.5 72B Instruct67.92
5Qwen 2.5 72B67.38
6QwQ 32B-Preview67.01
7Qwen 2.5 32B66.73
8Gemma 3 27B (IT)66.48
9Llama 3.3 70B Instruct66.4
10Qwen 2 72B66.02
11Qwen 2 72B Instruct65.87
12MSH-v1-Bielik-v2.3-Instruct-MedIT-merge65.8
13Bielik-11B-v2.3-Instruct65.71
14Bielik-11B-v2.2-Instruct65.57
15Llama 3.1 70B Instruct65.49

Interactive version: theaggregate.ai/benchmark?slug=open-pl-llm-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.