BenCzechMark: leaderboard

Multitask Czech language benchmark evaluating LLMs across machine learning, chat, and familiarity tasks. 64+ models tested with statistical significance-based scoring.

Metric: Average Score (%). Source: huggingface.co. Status: saturation imminent. 74 models tracked.

Top models

#ModelScore
1DeepSeek V3 (0324)83.39
2Llama 3.1 405B Instruct83.19
3DeepSeek V4 Flash (0731)81.36
4DeepSeek R1 052881.06
5c4ai-command-a-03-202578.09
6Qwen 2.5 72B77.91
7Mistral Large 2 (Nov) Instruct (2411)74.24
8Gemma 4 31B (IT)72.45
9Llama 3.3 70B Instruct71.95
10MiniMax M1 80k71.78
11Qwen 2.5 72B Instruct71.26
12Qwen 3.6 27B70.57
13Llama 4 Scout Instruct69.79
14Llama 3.1 70B Instruct69.51
15Qwen 2.5 32B Instruct66.04

Interactive version: theaggregate.ai/benchmark?slug=benczechmark · How It Works · Data refreshed daily, snapshot 2026-09-05.