HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Professional Psychology: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.793.1
2GPT-5.4 (2026-03-05)91.5
3Claude Sonnet 4.690.8
4GPT-4.1 (2025-04-14)89
5Gemini 2.5 Flash (Thinking)88.7
6Gemma 4 31B (IT)86
7Qwen 3.5 397B A17B85.9
8Mistral Large 385.9
9Claude Haiku 4.5 (20251001)84.7
10Qwen 3 Next 80B A3B Instruct82.5
11Qwen 2.5 72B Instruct81.7
12Llama 4 Maverick Instruct FP881.7
13DeepSeek V3.181.5
14GPT-4.1 Mini80.5
15Gemini 2.5 Flash Lite (Thinking)80.5

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-professional-psychology · How It Works · Data refreshed daily, snapshot 2026-09-19.