HELM Arabic - Alghafa - Multiple Choice Facts Truefalse Balanced Task: leaderboard

Metric: Exact Match (%). Source: crfm.stanford.edu. 41 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct97.33
2AceGPT-v2-32B-Chat97.33
3jais-adapted-70B-chat96
4Qwen 2.5 72B Instruct94.67
5Llama 4 Scout Instruct94.67
6Qwen 3.5 397B A17B94.67
7jais-family-30B-16k-chat94.67
8GPT-4.1 (2025-04-14)93.33
9GPT-5.4 (2026-03-05)93.33
10Gemma 4 31B (IT)92
11DeepSeek V3.192
12Mistral Large 2 (Nov) Instruct (2411)92
13Gemini 2.5 Flash Lite (Thinking)92
14Claude Opus 4.790.67
15Claude Haiku 4.5 (20251001)90.67

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-alghafa-multiple-choice-facts-truefalse-balanced-task · How It Works · Data refreshed daily, snapshot 2026-09-19.