HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Moral Scenarios: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.782.23
2Gemma 4 31B (IT)71.62
3Claude Sonnet 4.668.27
4Llama 3.3 70B Instruct65.59
5Llama 4 Maverick Instruct FP864.13
6Gemini 2.5 Flash (Thinking)63.8
7GPT-5.4 (2026-03-05)61.79
8Qwen 3.5 397B A17B58.44
9Claude Haiku 4.5 (20251001)58.32
10Llama 4 Scout Instruct56.09
11Qwen 3 Next 80B A3B Instruct56.09
12AceGPT-v2-70B-Chat54.64
13GPT-4.1 (2025-04-14)54.53
14Qwen 2.5 72B Instruct53.18
15Mistral Large 351.51

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-moral-scenarios · How It Works · Data refreshed daily, snapshot 2026-09-19.