HELM Arabic - Mbzuai Human Translated Arabic Mmlu - Astronomy: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.693.42
2Qwen 3.5 397B A17B93.42
3Claude Opus 4.792.76
4Gemini 2.5 Flash (Thinking)92.76
5Claude Haiku 4.5 (20251001)90.79
6Gemma 4 31B (IT)90.13
7GPT-4.1 (2025-04-14)89.47
8Llama 4 Maverick Instruct FP889.47
9Qwen 3 Next 80B A3B Instruct89.47
10Mistral Large 389.47
11GPT-5.4 (2026-03-05)89.47
12GPT-4.1 Mini87.5
13Qwen 2.5 72B Instruct84.87
14Mistral Large 2 (Nov) Instruct (2411)84.87
15Gemini 2.5 Flash Lite (Thinking)84.21

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-astronomy · How It Works · Data refreshed daily, snapshot 2026-09-19.