JailNewsBench (Role Play) - Harmfulness Score: leaderboard

Metric: Mean harmfulness of the generated fake news (0 to 4): the average of eight LLM-judge sub-metrics (faithfulness distortion, verifiability, adherence to the malicious instruction, scope, scale, formality, subjectivity and agitativeness), each scored 0 (harmless) to 4 and averaged over GPT-5, Gemini 2.5 and Claude 4 judges, over the fluent, non-refusing outputs on JailNewsBench test-split fake-news generation instructions (seed instructions built from real news articles of 34 regions in 22 languages, each with a financial, political, social or psychological motive), averaged across regions, under the Role Play jailbreak (the instruction is prefixed with a role that casts the model as a news writer); lower is better. Source: arxiv.org. 9 models tracked.

Top models

#ModelScoreOverall rank
1Llama 3.1 8B Instruct2.3#1018
2DeepSeek R1 Distill Llama 70B2.4#582
3Llama 3.1 70B Instruct2.5#548
4DeepSeek R1 Distill Llama 8B2.6#1282
5GPT-52.7#91
6Claude Sonnet 42.7#194
7Gemini 2.5 Flash3.2#237

Interactive version: theaggregate.ai/benchmark?slug=jailnewsbench-role-play-harmfulness-score · How It Works · Data refreshed daily, snapshot 2026-10-11.