IslamicLegalBench (Zero-Shot) - Hallucination Rate: leaderboard
Metric: Hallucination rate (%): share of answers that are incorrect or partly correct with fabricated or misattributed content (refusals not counted), the unweighted mean of the 13 task rates, on IslamicLegalBench's 718 expert-built Islamic-law questions in 13 task types over classical fiqh texts of several schools (metadata and concept recall, source attribution, condition enumeration, statute synthesis, 'illah identification, analogy and maxim mapping), answers graded against the reference for every core element by an o3 judge (correct, partially correct, incorrect or refusal), temperature 0, zero-shot prompting; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 | 20.53 | #91 |
| 2 | Claude Sonnet 4.5 | 23.33 | #138 |
| 3 | Grok 4 | 23.95 | #169 |
| 4 | Gemini 2.5 Pro | 27.9 | #145 |
| 5 | Llama 4 Maverick Instruct FP8 | 30.71 | #403 |
| 6 | DeepSeek R1 | 35.91 | #245 |
| 7 | Qwen 3 235B A22B 2507 Instruct | 44.17 | #291 |
| 8 | GPT-OSS-120B | 55.18 | #330 |
| 9 | Llama 3.1 8B Instruct | 61.67 | #1018 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=islamiclegalbench-zero-shot-hallucination-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.