Open Persian LLM Leaderboard — leaderboard

Persian language LLM evaluation: benchmarks models on Part Multiple Choice, ARC Easy/Challenge, MMLU Pro, and AUT Multiple Choice in Persian.

Metric: Average Score (%). Source: huggingface.co. Status: saturation imminent. 31 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct70.36
2Gemma 3 27B (IT)67.85
3Gemma 2 27B (IT)65.46
4QwQ 32B-Preview64.78
5Qwen 2.5 32B Instruct64.46
6Gemma 2 9B (IT)62.89
7aya-expanse-32B61.94
8Gemma 3 12B (IT)60.9
9aya-23-35B56.78
10aya-expanse-8B53.68
11Qwen 2.5 7B Instruct51.9
12Qwen 2 7B Instruct51.44
13Llama 3.1 8B Instruct50.14
14aya-23-8B49.84
15Gemma 3 4B (IT)49.06

Interactive version: theaggregate.ai/benchmark?slug=open-persian-llm-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.