HELM Arabic - Mbzuai Human Translated Arabic Mmlu - High School Macroeconomics: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.793.4
2GPT-5.4 (2026-03-05)91.9
3Claude Sonnet 4.691.5
4Gemini 2.5 Flash (Thinking)90.5
5GPT-4.1 (2025-04-14)89.3
6Mistral Large 387.7
7Claude Haiku 4.5 (20251001)87.1
8Qwen 3.5 397B A17B87.1
9Gemma 4 31B (IT)86.5
10Qwen 3 Next 80B A3B Instruct84.6
11DeepSeek V3.183
12Gemini 2.5 Flash Lite (Thinking)82.5
13Llama 4 Maverick Instruct FP882.2
14Qwen 2.5 72B Instruct81.8
15GPT-4.1 Mini81.4

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-high-school-macroeconomics · How It Works · Data refreshed daily, snapshot 2026-09-19.