NESSiE: leaderboard
Metric: Safe & Helpful rate (%): share of NESSiE test cases solved on both paired user prompts under one system prompt, withholding the protected keyword or secret on the unauthorized request and giving the required keyword on the authorized one, judged by keyword matching; 93 system-user combinations over 41 test cases in six suites (RULeS non-game, RULeS reformulated, agentic, generated, skills, multi-turn), three seeds and three runs; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.5 Pro | 95.2 | #145 |
| 2 | GPT-OSS-20B | 93.4 | #499 |
| 3 | GPT-5 | 92.6 | #91 |
| 4 | GPT-5.1 | 92.3 | #131 |
| 5 | Grok 4 | 92 | #169 |
| 6 | Gemini 3 Flash (Preview) | 88.9 | #78 |
| 7 | GPT-5.2 | 87.5 | #105 |
| 8 | GPT-4.1 Mini | 87.2 | #346 |
| 9 | Claude Opus 4.5 | 82.6 | #79 |
| 10 | Gemini 2.5 Flash | 81.8 | #237 |
| 11 | Claude Sonnet 4.5 | 81.8 | #138 |
| 12 | Claude Sonnet 4 | 80.3 | #194 |
| 13 | Qwen 3 VL 32B Instruct | 62.4 | #276 |
| 14 | Qwen 3 VL 8B Instruct | 55.8 | #401 |
| 15 | Mistral Instruct v0.1 7B | 29.1 | #1429 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=nessie · How It Works · Data refreshed daily, snapshot 2026-10-11.