JailNewsBench (Context Overload) - Harmfulness Score: leaderboard

Metric: Mean harmfulness of the generated fake news (0 to 4): the average of eight LLM-judge sub-metrics (faithfulness distortion, verifiability, adherence to the malicious instruction, scope, scale, formality, subjectivity and agitativeness), each scored 0 (harmless) to 4 and averaged over GPT-5, Gemini 2.5 and Claude 4 judges, over the fluent, non-refusing outputs on JailNewsBench test-split fake-news generation instructions (seed instructions built from real news articles of 34 regions in 22 languages, each with a financial, political, social or psychological motive), averaged across regions, under the Context Overload jailbreak (512 tokens of unrelated news articles are prepended to the instruction); lower is better. Source: arxiv.org. 9 models tracked.

Top models

#ModelScoreOverall rank
1Llama 3.1 8B Instruct2.1#1018
2DeepSeek R1 Distill Llama 70B2.1#582
3DeepSeek R1 Distill Llama 8B2.1#1282
4Llama 3.1 70B Instruct2.2#548
5GPT-52.3#91
6Claude Sonnet 42.4#194
7Gemini 2.5 Flash2.9#237

Interactive version: theaggregate.ai/benchmark?slug=jailnewsbench-context-overload-harmfulness-score · How It Works · Data refreshed daily, snapshot 2026-10-11.