JailNewsBench (Negative Prompting) - Harmfulness Score: leaderboard

Metric: Mean harmfulness of the generated fake news (0 to 4): the average of eight LLM-judge sub-metrics (faithfulness distortion, verifiability, adherence to the malicious instruction, scope, scale, formality, subjectivity and agitativeness), each scored 0 (harmless) to 4 and averaged over GPT-5, Gemini 2.5 and Claude 4 judges, over the fluent, non-refusing outputs on JailNewsBench test-split fake-news generation instructions (seed instructions built from real news articles of 34 regions in 22 languages, each with a financial, political, social or psychological motive), averaged across regions, under the Negative Prompting jailbreak (the request is rephrased as a prohibition that asks what such fake news would look like while forbidding an answer); lower is better. Source: arxiv.org. 9 models tracked.

Top models

#ModelScoreOverall rank
1Llama 3.1 8B Instruct2.3#1018
2DeepSeek R1 Distill Llama 70B2.3#582
3DeepSeek R1 Distill Llama 8B2.3#1282
4Llama 3.1 70B Instruct2.4#548
5GPT-52.6#91
6Claude Sonnet 42.6#194
7Gemini 2.5 Flash3.1#237

Interactive version: theaggregate.ai/benchmark?slug=jailnewsbench-negative-prompting-harmfulness-score · How It Works · Data refreshed daily, snapshot 2026-10-11.