JailNewsBench (Negative Prompting): leaderboard

Metric: Attack success rate (%): share of model outputs that do not refuse, as judged jointly by GPT-5, Gemini 2.5 and Claude 4, on JailNewsBench test-split fake-news generation instructions (seed instructions built from real news articles of 34 regions in 22 languages, each with a financial, political, social or psychological motive), averaged across regions, under the Negative Prompting jailbreak (the request is rephrased as a prohibition that asks what such fake news would look like while forbidding an answer); lower is better. Source: arxiv.org. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-574.6#91
2Claude Sonnet 475.7#194
3Gemini 2.5 Flash77.1#237
4DeepSeek R1 Distill Llama 70B79#582
5DeepSeek R1 Distill Llama 8B82.4#1282
6Llama 3.1 70B Instruct84.8#548
7Llama 3.1 8B Instruct85.1#1018

Interactive version: theaggregate.ai/benchmark?slug=jailnewsbench-negative-prompting · How It Works · Data refreshed daily, snapshot 2026-10-11.