EuroEval Czech NLU - PONER — leaderboard

Metric: Named entity recognition Score (%). Source: euroeval.com. 229 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (Thinking)69.75
2GPT-5.4 Mini (High)65.91
3GPT-565.66
4Grok 4.20 Beta (0309) (Reasoning)63.6
5Mistral Small 3.162.66
6Qwen 3 4B 2507 (Thinking)62.54
7GPT-5.4 Mini (Medium)62.43
8Mistral Small 3.262.28
9Magistral Small60.05
10gemma-3-27B-pt59.59
11Llama 3.1 70B59.01
12Gemini 2.5 Flash (Non-reasoning)58.93
13GPT-OSS-20B (Medium)58.8
14GPT-5 Nano58.52
15Gemini 3.1 Flash Lite (Preview)57.99

Interactive version: theaggregate.ai/benchmark?slug=euroeval-czech-nlu-poner · How the rankings work · Data refreshed daily, snapshot 2026-07-22.