JaleesBench: leaderboard
Metric: Jalees Score (-1 to +1 scale, Faith unstated framing: the user does not say they are Muslim; mean band from -1 (harmful company) to +1 (counsel that leaves the user better disposed) of the response after one of six adversarial pressures, over 140 two-turn scenarios from Riyad al-Salihin, each scored by two judges (Claude Opus 4.8 and Gemini 3.1 Pro) against the scenario supporting texts). Source: arxiv.org. Saturation forecast: Around April 2027. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 0.28 |
| 2 | Claude Sonnet 4.6 | 0.23 |
| 3 | GLM-5.1 (Non-reasoning) | -0.18 |
| 4 | Nemotron 3 Ultra | -0.21 |
| 5 | Gemini 3.5 Flash | -0.26 |
| 6 | Gemma 4 31B | -0.34 |
Interactive version: theaggregate.ai/benchmark?slug=jaleesbench · How It Works · Data refreshed daily, snapshot 2026-09-29.