HELM Arabic - Alghafa - MCQ Exams Test AR: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct64.99
2Claude Opus 4.762.66
3GPT-4.1 (2025-04-14)61.58
4Claude Sonnet 4.661.4
5GPT-5.4 (2026-03-05)61.4
6Llama 4 Maverick Instruct FP860.86
7Gemini 2.5 Flash (Thinking)60.32
8Llama 4 Scout Instruct59.96
9Qwen 3.5 397B A17B59.78
10Mistral Large 359.78
11Gemini 2.5 Flash Lite (Thinking)59.61
12Qwen 2.5 72B Instruct59.43
13Gemma 4 31B (IT)59.25
14Claude Haiku 4.5 (20251001)59.07
15Qwen 3 Next 80B A3B Instruct58.89

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-alghafa-mcq-exams-test-ar · How It Works · Data refreshed daily, snapshot 2026-09-19.