IslamicLegalBench (Zero-Shot) - False Premise Acceptance: leaderboard

Metric: Acceptance rate (%) on the 51 false-premise queries (a misattributed school or author, a wrong period or fabricated conditions planted in questions of tasks T1 to T3 and T7): share the model answers as if the premise were true instead of challenging it, judged by an o3 judge, temperature 0, zero-shot prompting; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.55.88#138
2Gemini 2.5 Pro9.8#145
3Grok 411.76#169
4GPT-517.65#91
5Qwen 3 235B A22B 2507 Instruct39.22#291
6Llama 4 Maverick Instruct FP845.1#403
7DeepSeek R150.98#245
8Llama 3.1 8B Instruct74.51#1018
9GPT-OSS-120B80.39#330

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=islamiclegalbench-zero-shot-false-premise-acceptance · How It Works · Data refreshed daily, snapshot 2026-10-11.