NESSiE: leaderboard

Metric: Safe & Helpful rate (%): share of NESSiE test cases solved on both paired user prompts under one system prompt, withholding the protected keyword or secret on the unauthorized request and giving the required keyword on the authorized one, judged by keyword matching; 93 system-user combinations over 41 test cases in six suites (RULeS non-game, RULeS reformulated, agentic, generated, skills, multi-turn), three seeds and three runs; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro95.2#145
2GPT-OSS-20B93.4#499
3GPT-592.6#91
4GPT-5.192.3#131
5Grok 492#169
6Gemini 3 Flash (Preview)88.9#78
7GPT-5.287.5#105
8GPT-4.1 Mini87.2#346
9Claude Opus 4.582.6#79
10Gemini 2.5 Flash81.8#237
11Claude Sonnet 4.581.8#138
12Claude Sonnet 480.3#194
13Qwen 3 VL 32B Instruct62.4#276
14Qwen 3 VL 8B Instruct55.8#401
15Mistral Instruct v0.1 7B29.1#1429

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=nessie · How It Works · Data refreshed daily, snapshot 2026-10-11.