HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Formal Logic: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.784.92
2GPT-5.4 (2026-03-05)73.02
3Claude Haiku 4.5 (20251001)72.22
4Gemini 2.5 Flash (Thinking)72.22
5Claude Sonnet 4.670.63
6Qwen 3.5 397B A17B70.63
7Gemma 4 31B (IT)69.84
8Qwen 3 Next 80B A3B Instruct69.05
9GPT-4.1 (2025-04-14)66.67
10Mistral Large 365.08
11Llama 4 Maverick Instruct FP864.29
12GPT-4.1 Mini61.11
13Qwen 2.5 72B Instruct61.11
14DeepSeek V3.159.52
15Gemini 2.5 Flash Lite (Thinking)59.52

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-formal-logic · How It Works · Data refreshed daily, snapshot 2026-09-19.