HELM Arabic - Arabic Mmlu - Law (Professional): leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.787.26
2GPT-5.4 (2026-03-05)86.31
3Claude Sonnet 4.683.76
4Llama 4 Maverick Instruct FP882.48
5Gemma 4 31B (IT)80.57
6Qwen 2.5 72B Instruct78.03
7Llama 4 Scout Instruct77.39
8GPT-4.1 (2025-04-14)76.75
9Mistral Large 376.11
10DeepSeek V3.174.84
11Gemini 2.5 Flash (Thinking)71.97
12GPT-4.1 Mini70.38
13Llama 3.3 70B Instruct70.06
14Claude Haiku 4.5 (20251001)69.11
15AceGPT-v2-32B-Chat67.83

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-arabic-mmlu-law-professional · How It Works · Data refreshed daily, snapshot 2026-09-19.