HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Conceptual Physics: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.793.19
2Claude Sonnet 4.692.77
3Qwen 3.5 397B A17B88.09
4Gemini 2.5 Flash (Thinking)86.81
5Gemma 4 31B (IT)85.53
6Claude Haiku 4.5 (20251001)84.26
7GPT-5.4 (2026-03-05)83.83
8Llama 4 Maverick Instruct FP882.13
9GPT-4.1 (2025-04-14)81.28
10Qwen 3 Next 80B A3B Instruct81.28
11DeepSeek V3.181.28
12Qwen 3.5 9B80.43
13Mistral Large 380.43
14Gemini 2.5 Flash Lite (Thinking)78.72
15Llama 4 Scout Instruct75.32

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-conceptual-physics · How It Works · Data refreshed daily, snapshot 2026-09-19.