HELM Arabic - Arabic Exams - Social: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct55.51
2Qwen 2.5 72B Instruct51.1
3Claude Haiku 4.5 (20251001)50.74
4Gemini 2.5 Flash (Thinking)50.74
5Llama 4 Maverick Instruct FP850.37
6Llama 4 Scout Instruct50
7GPT-4.1 (2025-04-14)50
8Claude Opus 4.750
9Mistral Large 349.63
10Gemini 2.5 Flash Lite (Thinking)49.63
11GPT-5.4 (2026-03-05)49.26
12DeepSeek V3.148.9
13Claude Sonnet 4.648.53
14GPT-4.1 Mini47.79
15Gemma 4 31B (IT)47.79

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-arabic-exams-social · How It Works · Data refreshed daily, snapshot 2026-09-19.