HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Sociology: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.792.04
2Mistral Large 389.05
3GPT-5.4 (2026-03-05)89.05
4Gemini 2.5 Flash (Thinking)88.56
5Qwen 3.5 397B A17B88.06
6Claude Sonnet 4.687.56
7Claude Haiku 4.5 (20251001)87.06
8GPT-4.1 (2025-04-14)85.57
9Gemma 4 31B (IT)84.58
10Qwen 2.5 72B Instruct84.08
11Llama 4 Maverick Instruct FP884.08
12DeepSeek V3.183.58
13Gemini 2.5 Flash Lite (Thinking)83.08
14Llama 3.3 70B Instruct81.09
15Llama 4 Scout Instruct80.1

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-sociology · How It Works · Data refreshed daily, snapshot 2026-09-19.