IslamicLegalBench (Few-Shot) - False Premise Acceptance: leaderboard

Metric: Acceptance rate (%) on the 51 false-premise queries (a misattributed school or author, a wrong period or fabricated conditions planted in questions of tasks T1 to T3 and T7): share the model answers as if the premise were true instead of challenging it, judged by an o3 judge, temperature 0, few-shot prompting with format and reasoning exemplars; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro7.84#145
2Claude Sonnet 4.57.84#138
3GPT-517.65#91
4Grok 425.49#169
5Llama 4 Maverick Instruct FP841.18#403
6Qwen 3 235B A22B 2507 Instruct47.06#291
7DeepSeek R160.78#245
8Llama 3.1 8B Instruct72.55#1018
9GPT-OSS-120B86.27#330

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=islamiclegalbench-few-shot-false-premise-acceptance · How It Works · Data refreshed daily, snapshot 2026-10-11.