HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Econometrics: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.780.7
2GPT-5.4 (2026-03-05)78.95
3Claude Sonnet 4.673.68
4Qwen 3.5 397B A17B72.81
5Gemma 4 31B (IT)67.54
6Claude Haiku 4.5 (20251001)64.91
7Llama 4 Maverick Instruct FP864.04
8DeepSeek V3.162.28
9Mistral Large 362.28
10Gemini 2.5 Flash (Thinking)62.28
11Qwen 3 Next 80B A3B Instruct61.4
12GPT-4.1 (2025-04-14)60.53
13Llama 4 Scout Instruct58.77
14GPT-4.1 Mini55.26
15Llama 3.3 70B Instruct53.51

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-econometrics · How It Works · Data refreshed daily, snapshot 2026-09-19.