EuroEval Catalan NLU - MultiWikiQA CA — leaderboard
Metric: Reading comprehension Score (%). Source: euroeval.com. 212 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Ministral 3 3B | 78.35 |
| 2 | SOLAR-10.7B-v1.0 | 78.2 |
| 3 | Mistral Small 3.2 | 77.99 |
| 4 | gemma-3-12B-pt | 77.7 |
| 5 | Olmo 3.1 32B Instruct | 77.23 |
| 6 | Ministral-3-14B-Reasoning-2512 | 76.9 |
| 7 | Qwen 3 14B (Non-reasoning) | 76.65 |
| 8 | Bielik-11B-v2.3-Instruct | 76.56 |
| 9 | Mistral Small 3.1 | 75.96 |
| 10 | Llama 3.1 70B | 75.86 |
| 11 | Ministral 3 14B | 74.48 |
| 12 | Ministral-3-3B-Reasoning-2512 | 74.41 |
| 13 | Qwen 3 8B (Non-reasoning) | 74.17 |
| 14 | gemma-3-27B-pt | 74.03 |
| 15 | Llama 3.1 8B | 73.75 |
Interactive version: theaggregate.ai/benchmark?slug=euroeval-catalan-nlu-multiwikiqa-ca · How the rankings work · Data refreshed daily, snapshot 2026-07-22.