HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Security Studies: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.680.82
2GPT-5.4 (2026-03-05)80.82
3Gemma 4 31B (IT)80.41
4Claude Opus 4.780
5Qwen 3.5 397B A17B78.37
6Llama 4 Maverick Instruct FP876.73
7Mistral Large 375.92
8Mistral Large 2 (Nov) Instruct (2411)75.92
9Claude Haiku 4.5 (20251001)75.51
10Gemini 2.5 Flash (Thinking)75.51
11DeepSeek V3.175.1
12Llama 3.3 70B Instruct73.88
13GPT-4.1 (2025-04-14)73.88
14Llama 4 Scout Instruct73.47
15Gemini 2.5 Flash Lite (Thinking)73.47

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-security-studies · How It Works · Data refreshed daily, snapshot 2026-09-19.