HELM Arabic - Alghafa - Multiple Choice Rating Sentiment Task: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.663.1
2GPT-4.1 (2025-04-14)63
3Gemma 4 31B (IT)60.2
4Claude Opus 4.760.2
5Llama 4 Maverick Instruct FP859.6
6Llama 3.3 70B Instruct57.9
7jais-adapted-7B-chat56.9
8Llama 4 Scout Instruct56.4
9AceGPT-v2-8B-Chat56.4
10jais-adapted-70B-chat56.3
11Qwen 3.5 397B A17B56.1
12Gemini 2.5 Flash (Thinking)56
13AceGPT-v2-70B-Chat56
14Claude Haiku 4.5 (20251001)55.8
15AceGPT-v2-32B-Chat55.7

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-alghafa-multiple-choice-rating-sentiment-task · How It Works · Data refreshed daily, snapshot 2026-09-19.