HELM Arabic - Arabic Mmlu - Political Science (University): leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.685.24
2Claude Opus 4.785.24
3GPT-5.4 (2026-03-05)82.86
4Gemma 4 31B (IT)81.43
5Gemini 2.5 Flash (Thinking)80.48
6Claude Haiku 4.5 (20251001)79.05
7Qwen 3 Next 80B A3B Instruct78.1
8Gemini 2.5 Flash Lite (Thinking)77.62
9Llama 4 Maverick Instruct FP877.14
10GPT-4.1 (2025-04-14)76.67
11Qwen 3.5 397B A17B75.71
12Qwen 2.5 72B Instruct73.81
13Llama 4 Scout Instruct73.33
14Llama 3.3 70B Instruct72.86
15AceGPT-v2-70B-Chat71.43

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-arabic-mmlu-political-science-university · How It Works · Data refreshed daily, snapshot 2026-09-19.