HELM Arabic - Alghafa - Multiple Choice Sentiment Task: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Haiku 4.5 (20251001)46.1
2Qwen 3.5 397B A17B45.5
3GPT-4.1 (2025-04-14)43.9
4Llama 4 Scout Instruct43.7
5DeepSeek V3.143.7
6Gemini 2.5 Flash (Thinking)43.7
7Llama 3.3 70B Instruct43.3
8GPT-5.4 (2026-03-05)43.1
9Qwen 2.5 72B Instruct43
10Claude Opus 4.743
11Llama 4 Maverick Instruct FP842.7
12Gemini 2.5 Flash Lite (Thinking)42.2
13Claude Sonnet 4.642.1
14AceGPT-v2-32B-Chat42.1
15Mistral Large 341.9

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-alghafa-multiple-choice-sentiment-task · How It Works · Data refreshed daily, snapshot 2026-09-19.