Turkish MMLU — leaderboard
Turkish-language MMLU evaluation with 6,200 multiple-choice questions testing knowledge across diverse academic subjects in Turkish.
Metric: Accuracy (%). Source: huggingface.co. Status: saturation imminent. 66 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 84.8 |
| 2 | Claude 3.5 Sonnet (20240620) | 84.4 |
| 3 | Gemini 1.5 Pro | 76.7 |
| 4 | Qwen 3 32B | 76 |
| 5 | Gemma 3 27B | 75.1 |
| 6 | Gemma 2 27B | 72.1 |
| 7 | Qwen 3 14B | 71.7 |
| 8 | Gemma 3 12B | 70.7 |
| 9 | aya-expanse-32B | 70.7 |
| 10 | Llama 3.1 70B | 70.4 |
| 11 | mistral-small-latest | 67 |
| 12 | Gemma 3n E4B | 63.7 |
| 13 | Gemma 2 2B | 48.4 |
| 14 | Gemma 3 4B | 42.7 |
| 15 | Gemma 3 1B | 42.7 |
Interactive version: theaggregate.ai/benchmark?slug=turkish-mmlu · How the rankings work · Data refreshed daily, snapshot 2026-07-22.