Open Arabic LLM - Arabic MMLU HT Formal Logic — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 163 models tracked.

Top models

#ModelScore
1Qwen3-8B-Base61.9
2Ultiima-72B61.11
3Qwen2.5-32B-Instruct-CFT59.52
4calme-2.1-qwen2.5-72B59.52
5calme-2.2-qwen2.5-72B59.52
6Rombos-LLM-V2.5-Qwen-72B59.52
7Gemma 3 27B (IT)58.73
8Llama 3.3 70B Instruct58.73
9lambda-qwen2.5-14B-dpo-test58.73
10Awqward2.5-32B-Instruct57.94
11tempmotacilla-cinerea-030857.94
12lambda-qwen2.5-32B-dpo-test57.94
13Qwen 2.5 32B Instruct57.14
14Qwen2.5-32B-Instruct-abliterated-v257.14
15Qwen 2.5 72B Instruct56.35

Interactive version: theaggregate.ai/benchmark?slug=open-arabic-llm-arabic-mmlu-ht-formal-logic · How the rankings work · Data refreshed daily, snapshot 2026-07-22.