EuroEval Croatian NLU - Multi Wiki QA HR: leaderboard
Metric: Reading comprehension Score (%). Source: euroeval.com. 211 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Ministral-3-14B-Reasoning-2512 | 73.9 |
| 2 | Olmo-3.1-32B-Instruct-SFT | 70.12 |
| 3 | Ministral 3 14B | 69.62 |
| 4 | SOLAR-10.7B-v1.0 | 68.75 |
| 5 | Mistral Small 3.2 | 68.62 |
| 6 | Ministral-3-14B-Instruct-2512 | 67.67 |
| 7 | Gemini 3 Pro (Preview) | 67.64 |
| 8 | GPT-5.2 | 67.02 |
| 9 | Llama 3.1 70B | 66.7 |
| 10 | Ministral-3-3B-Reasoning-2512 | 66.43 |
| 11 | GPT-5.4 Mini (Medium) | 66.25 |
| 12 | Gemini 2.5 Flash (Thinking) | 65.65 |
| 13 | gemma-4-E2B-it | 65.47 |
| 14 | GPT-5 Nano | 65.17 |
| 15 | Olmo-3.1-32B-Instruct-DPO | 65.08 |
Interactive version: theaggregate.ai/benchmark?slug=euroeval-croatian-nlu-multi-wiki-qa-hr · How It Works · Data refreshed daily, snapshot 2026-09-05.