IslamicLegalBench - High Complexity — leaderboard

Metric: Weighted Accuracy (%). Source: huggingface.co. 12 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.575.5
2Gemini 2.5 Pro73.12
3GPT-572.99
4Grok 472.88
5Qwen 3 235B A22B 2507 Instruct67.65
6DeepSeek R166.51
7Llama 4 Maverick Instruct FP855.14
8Llama 3.1 8B Instruct52.99
9GPT-OSS-120B50.7

Interactive version: theaggregate.ai/benchmark?slug=islamiclegalbench-high-complexity · How the rankings work · Data refreshed daily, snapshot 2026-07-22.