IslamicMMLU - Hadith Cloze Completion: leaderboard

Metric: Accuracy (%) on IslamicMMLU's 1,000 Hadith track questions on filling in a missing keyword of a hadith; four-option multiple choice in Arabic, zero-shot, temperature 0, answer letter only (10 new tokens), via each model's API; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 26 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash95.8#93
2Gemini 3 Pro95.7#77
3Claude Opus 4.595.6#79
4GPT-595.2#91
5Gemini 2.5 Pro95#145
6GPT-5.194#131
7GPT-5.292.1#105
8Claude Sonnet 4.591#138
9Claude 3.7 Sonnet90.2#241
10Gemini 2.5 Flash85.3#237
11DeepSeek V3.184.6#260
12GPT-4.183.5#240
13DeepSeek V3.283.3#198
14GPT-4o81.8#333
15Grok 4.1 Fast78.1#208

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=islamicmmlu-hadith-cloze-completion · How It Works · Data refreshed daily, snapshot 2026-10-11.