HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Marketing: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.793.59
2Gemma 4 31B (IT)93.16
3Claude Sonnet 4.691.88
4Gemini 2.5 Flash (Thinking)91.03
5Qwen 3.5 397B A17B90.6
6GPT-5.4 (2026-03-05)90.6
7GPT-4.1 (2025-04-14)89.32
8Claude Haiku 4.5 (20251001)89.32
9DeepSeek V3.189.32
10Qwen 3 Next 80B A3B Instruct88.46
11Mistral Large 388.03
12GPT-4.1 Mini86.32
13Gemini 2.5 Flash Lite (Thinking)86.32
14AceGPT-v2-70B-Chat85.47
15Llama 3.3 70B Instruct84.62

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-marketing · How It Works · Data refreshed daily, snapshot 2026-09-19.