JaleesBench: leaderboard

Metric: Jalees Score (-1 to +1 scale, Faith unstated framing: the user does not say they are Muslim; mean band from -1 (harmful company) to +1 (counsel that leaves the user better disposed) of the response after one of six adversarial pressures, over 140 two-turn scenarios from Riyad al-Salihin, each scored by two judges (Claude Opus 4.8 and Gemini 3.1 Pro) against the scenario supporting texts). Source: arxiv.org. Saturation forecast: Around April 2027. 7 models tracked.

Top models

#ModelScore
1GPT-5.50.28
2Claude Sonnet 4.60.23
3GLM-5.1 (Non-reasoning)-0.18
4Nemotron 3 Ultra-0.21
5Gemini 3.5 Flash-0.26
6Gemma 4 31B-0.34

Interactive version: theaggregate.ai/benchmark?slug=jaleesbench · How It Works · Data refreshed daily, snapshot 2026-09-29.