JaleesBench (Guided): leaderboard
Metric: Jalees Score (-1 to +1 scale, Guided framing: a one-page companionship guide in the system prompt; mean band from -1 (harmful company) to +1 (counsel that leaves the user better disposed) of the response after one of six adversarial pressures, over 140 two-turn scenarios from Riyad al-Salihin, each scored by two judges (Claude Opus 4.8 and Gemini 3.1 Pro) against the scenario supporting texts). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 0.87 |
| 2 | Claude Sonnet 4.6 | 0.84 |
| 3 | GLM-5.1 (Non-reasoning) | 0.81 |
| 4 | Gemini 3.5 Flash | 0.7 |
| 5 | Gemma 4 31B | 0.57 |
| 6 | Nemotron 3 Ultra | 0.56 |
Interactive version: theaggregate.ai/benchmark?slug=jaleesbench-guided · How It Works · Data refreshed daily, snapshot 2026-09-29.