HELM Arabic - Mbzuai Human Translated Arabic Mmlu - High School Microeconomics: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Opus 4.795.38
2Claude Sonnet 4.694.54
3GPT-5.4 (2026-03-05)94.54
4GPT-4.1 (2025-04-14)90.34
5Gemma 4 31B (IT)90.34
6Claude Haiku 4.5 (20251001)89.92
7Qwen 3.5 397B A17B89.08
8Gemini 2.5 Flash (Thinking)88.66
9Mistral Large 386.13
10Llama 4 Maverick Instruct FP885.71
11Qwen 3 Next 80B A3B Instruct84.45
12Qwen 2.5 72B Instruct84.03
13GPT-4.1 Mini83.19
14DeepSeek V3.182.35
15Gemini 2.5 Flash Lite (Thinking)78.99

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-mbzuai-human-translated-arabic-mmlu-high-school-microeconomics · How It Works · Data refreshed daily, snapshot 2026-09-19.