HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Business Ethics: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.781
2GPT-5.4 (2026-03-05)81
3Qwen 3.5 397B A17B80
4Claude Sonnet 4.679
5GPT-4.1 (2025-04-14)79
6Claude Haiku 4.5 (20251001)77
7Mistral Large 377
8DeepSeek V3.176
9Gemini 2.5 Flash (Thinking)76
10GPT-4.1 Mini75
11Qwen 2.5 72B Instruct75
12Llama 4 Scout Instruct75
13Gemma 4 31B (IT)75
14Llama 4 Maverick Instruct FP875
15AceGPT-v2-70B-Chat75

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-business-ethics · How It Works · Data refreshed daily, snapshot 2026-09-19.