IslamicMMLU - Hadith Source Identification: leaderboard

Metric: Accuracy (%) on IslamicMMLU's 1,000 Hadith track questions on naming which of the six canonical collections a hadith (chain of narrators removed) comes from; four-option multiple choice in Arabic, zero-shot, temperature 0, answer letter only (10 new tokens), via each model's API; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 26 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash96.4#93
2Claude Opus 4.594.7#79
3Gemini 3 Pro93.1#77
4Gemini 2.5 Pro91.7#145
5Claude 3.7 Sonnet90.9#241
6Claude Sonnet 4.589.8#138
7GPT-5.289.6#105
8Gemini 2.5 Flash89.4#237
9GPT-587.6#91
10GPT-5.186#131
11GPT-4.182.1#240
12Llama 4 Maverick79.9#451
13GPT-4o79.5#333
14Claude Haiku 4.573.9#271
15DeepSeek V3.272#198

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=islamicmmlu-hadith-source-identification · How It Works · Data refreshed daily, snapshot 2026-10-11.