HELM Arabic - Alghafa - Multiple Choice Grounded Statement Soqal Task: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.696.67
2Gemma 4 31B (IT)96.67
3Mistral Large 396.67
4Gemini 2.5 Flash (Thinking)96
5Gemini 2.5 Flash Lite (Thinking)96
6GPT-4.1 (2025-04-14)95.33
7Llama 4 Maverick Instruct FP895.33
8DeepSeek V3.195.33
9Mistral Large 2 (Nov) Instruct (2411)95.33
10GPT-5.4 (2026-03-05)95.33
11Llama 3.3 70B Instruct94.67
12Qwen 2.5 72B Instruct94
13Llama 4 Scout Instruct94
14Claude Opus 4.794
15Claude Haiku 4.5 (20251001)94

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-alghafa-multiple-choice-grounded-statement-soqal-task · How It Works · Data refreshed daily, snapshot 2026-09-19.