BenCzechMark — leaderboard

Multitask Czech language benchmark evaluating LLMs across machine learning, chat, and familiarity tasks. 64+ models tested with statistical significance-based scoring.

Metric: Average Score (%). Source: huggingface.co. Status: saturated. 68 models tracked.

Top models

#ModelScore
1DeepSeek V3 (0324)86.4
2Llama 3.1 405B Instruct85.1
3DeepSeek R1 052884.82
4c4ai-command-a-03-202582.29
5Qwen 2.5 72B79.56
6MiniMax M1 80k76.36
7Llama 3.3 70B Instruct73.2
8Qwen 2.5 72B Instruct72.38
9Llama 3.1 70B Instruct72.3
10Llama 4 Scout Instruct71.33
11Llama 3.1 70B69.34
12Mistral Large 2 (Nov) Instruct (2411)68
13Qwen 2.5 32B Instruct67.2
14Gemma 4 31B (IT)66.12
15Qwen 2 72B Instruct64.8

Interactive version: theaggregate.ai/benchmark?slug=benczechmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.