IslamicMMLU - Hadith: leaderboard

Metric: Accuracy (%) on IslamicMMLU's 4,000 Hadith questions from the six canonical collections (source, cloze, chapter, authenticity grading); four-option multiple choice in Arabic, zero-shot, temperature 0, answer letter only (10 new tokens), via each model's API; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 26 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash93#93
2Gemini 3 Pro91.7#77
3Claude Opus 4.590.8#79
4Gemini 2.5 Pro89.7#145
5GPT-5.188.3#131
6GPT-588.2#91
7GPT-5.287.8#105
8Claude Sonnet 4.585.4#138
9Claude 3.7 Sonnet84.9#241
10Gemini 2.5 Flash83.8#237
11GPT-4o82#333
12GPT-4.181.3#240
13Llama 4 Maverick73.5#451
14Claude Haiku 4.573.1#271
15DeepSeek V3.271.1#198

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=islamicmmlu-hadith · How It Works · Data refreshed daily, snapshot 2026-10-11.