CPTU Bench — leaderboard

Polish language comprehension benchmark: evaluates LLMs on sentiment analysis, language understanding, phraseology, and tricky questions in Polish.

Metric: Average Score (1-5). Source: huggingface.co. Status: saturation imminent. 93 models tracked.

Top models

#ModelScore
1Gemini 2.0 Flash (001)4.29
2DeepSeek V3.24.14
3DeepSeek R14.14
4Gemini 2.0 Flash Lite (001)4.09
5DeepSeek V3.14.03
6DeepSeek V3 (0324)4.03
7DeepSeek V34.02
8Mistral Large 2 (Nov) Instruct (2411)4
9Kimi K2 09053.98
10Qwen 2.5 72B Instruct3.95
11Mistral Large 2 (Jul)3.93
12Llama 4 Maverick Instruct FP83.93
13Mistral Small 3.13.9
14Mistral Small 3.23.83
15GPT-OSS-120B3.82

Interactive version: theaggregate.ai/benchmark?slug=cptu-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.