INCLUDE-base-44 European Languages — leaderboard

European-language slice of INCLUDE-base-44, evaluating multilingual LLMs on knowledge- and reasoning-centric multiple-choice questions across 20 European languages.

Metric: Average Accuracy (%). Source: huggingface.co. Status: saturation imminent. 35 models tracked.

Top models

#ModelScore
1Qwen 3 14B63.03
2Gemma 3 12B (IT)62.55
3Qwen 2.5 14B Instruct60.56
4Qwen 2.5 14B58.31
5Qwen 3 8B58.25
6Phi-456.72
7Apertus-8B-Instruct-250956.16
8Llama 3.1 8B Instruct53.48
9Mistral Nemo Instruct (2407)52.2
10Qwen 2.5 7B51.78
11Qwen 2.5 7B Instruct51.41
12Mistral-Nemo-Base-240749.55
13Bielik-11B-v2.3-Instruct47.43
14Bielik-11B-v2.2-Instruct47.43
15Bielik-11B-v2.1-Instruct47.36

Interactive version: theaggregate.ai/benchmark?slug=include-base-44-european-languages · How the rankings work · Data refreshed daily, snapshot 2026-07-22.