HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Philosophy: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.789.07
2GPT-5.4 (2026-03-05)87.14
3Claude Sonnet 4.686.17
4Gemini 2.5 Flash (Thinking)81.35
5Claude Haiku 4.5 (20251001)80.71
6GPT-4.1 (2025-04-14)79.74
7Gemma 4 31B (IT)78.46
8Qwen 3.5 397B A17B78.14
9Qwen 3 Next 80B A3B Instruct76.85
10Mistral Large 376.85
11DeepSeek V3.174.92
12Llama 4 Maverick Instruct FP873.95
13Qwen 2.5 72B Instruct73.63
14Gemini 2.5 Flash Lite (Thinking)73.63
15AceGPT-v2-70B-Chat73.63

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-philosophy · How It Works · Data refreshed daily, snapshot 2026-09-19.