Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment Task — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 163 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct58.3
2free-evo-qwen72B-v0.8-re57.83
3Qwen 2 72B57.66
4aya-23-35B57.38
5Qwen 2.5 72B55.26
6Qwen 2.5 14B55.01
7aya-expanse-32B54.98
8Qwen 2.5 32B Instruct54.96
9Qwen2.5-32B-Instruct-CFT54.96
10Qwen 2.5 7B54.95
11lambda-qwen2.5-32B-dpo-test54.95
12Awqward2.5-32B-Instruct54.93
13Qwentile2.5-32B-Instruct54.93
14oxyge1-33B54.65
15Qwen2.5-32B-Instruct-abliterated-v254.61

Interactive version: theaggregate.ai/benchmark?slug=open-arabic-llm-alghafa-multiple-choice-rating-sentiment-task · How the rankings work · Data refreshed daily, snapshot 2026-07-22.