HELM Arabic Enterprise - Arabic Finance Mcq: leaderboard

Metric: EM. Source: crfm.stanford.edu. 35 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.695.65
2Claude Opus 4.795.65
3GPT-5.4 (2026-03-05)95.65
4DeepSeek V4 Pro91.3
5Gemma 4 31B (IT)86.96
6Claude Haiku 4.5 (20251001)86.96
7Mistral Large 386.96
8Gemini 2.5 Flash (Thinking)86.96
9Qwen 3.5 9B82.61
10Qwen 3.5 397B A17B82.61
11Mistral Large 2 (Nov) Instruct (2411)82.61
12AceGPT-v2-8B-Chat82.61
13Qwen3 235B A22B Instruct 2507 FP882.61
14Llama 3.3 70B Instruct78.26
15DeepSeek V3.178.26

Interactive version: theaggregate.ai/benchmark?slug=helm-arabic-enterprise-arabic-finance-mcq · How It Works · Data refreshed daily, snapshot 2026-09-08.