JailNewsBench (System Override) - Harmfulness Score: leaderboard

Metric: Mean harmfulness of the generated fake news (0 to 4): the average of eight LLM-judge sub-metrics (faithfulness distortion, verifiability, adherence to the malicious instruction, scope, scale, formality, subjectivity and agitativeness), each scored 0 (harmless) to 4 and averaged over GPT-5, Gemini 2.5 and Claude 4 judges, over the fluent, non-refusing outputs on JailNewsBench test-split fake-news generation instructions (seed instructions built from real news articles of 34 regions in 22 languages, each with a financial, political, social or psychological motive), averaged across regions, under the System Override jailbreak (the instruction is prefixed with a directive to ignore previous instructions and system prompts); lower is better. Source: arxiv.org. 9 models tracked.

Top models

#ModelScoreOverall rank
1Llama 3.1 8B Instruct2.5#1018
2DeepSeek R1 Distill Llama 8B2.5#1282
3DeepSeek R1 Distill Llama 70B2.8#582
4Llama 3.1 70B Instruct2.9#548
5Claude Sonnet 42.9#194
6GPT-53#91
7Gemini 2.5 Flash3.6#237

Interactive version: theaggregate.ai/benchmark?slug=jailnewsbench-system-override-harmfulness-score · How It Works · Data refreshed daily, snapshot 2026-10-11.