IslamicMMLU - Fiqh: leaderboard

Metric: Accuracy (%) on IslamicMMLU's 3,200 Fiqh knowledge questions from the four-school encyclopedia of al-Jaziri (the 800 madhab-bias questions are scored separately); four-option multiple choice in Arabic, zero-shot, temperature 0, answer letter only (10 new tokens), via each model's API; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 26 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash89.1#93
2Gemini 3 Pro87.9#77
3GPT-5.285.9#105
4Gemini 2.5 Pro85.5#145
5Claude Sonnet 4.584.9#138
6GPT-583.5#91
7Gemini 2.5 Flash83.1#237
8GPT-5.181.5#131
9Claude 3.7 Sonnet81.1#241
10GPT-4o78.5#333
11GPT-4.177.5#240
12DeepSeek V3.274.3#198
13Grok 4.1 Fast74#208
14Claude Opus 4.572.7#79
15Llama 4 Scout70.6#646

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=islamicmmlu-fiqh · How It Works · Data refreshed daily, snapshot 2026-10-11.