HELM Arabic - Mbzuai Human Translated Arabic Mmlu - International Law: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.792.56
2GPT-5.4 (2026-03-05)90.91
3GPT-4.1 (2025-04-14)89.26
4Gemini 2.5 Flash (Thinking)89.26
5Claude Sonnet 4.688.43
6Llama 4 Maverick Instruct FP887.6
7Qwen 3.5 397B A17B87.6
8Gemma 4 31B (IT)85.95
9Mistral Large 385.12
10Llama 3.3 70B Instruct83.47
11Qwen 3 Next 80B A3B Instruct83.47
12DeepSeek V3.183.47
13Claude Haiku 4.5 (20251001)82.64
14AceGPT-v2-70B-Chat81.82
15Qwen 2.5 72B Instruct80.17

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-international-law · How It Works · Data refreshed daily, snapshot 2026-09-19.