JailNewsBench (Context Overload) - Harmfulness Score: leaderboard
Metric: Mean harmfulness of the generated fake news (0 to 4): the average of eight LLM-judge sub-metrics (faithfulness distortion, verifiability, adherence to the malicious instruction, scope, scale, formality, subjectivity and agitativeness), each scored 0 (harmless) to 4 and averaged over GPT-5, Gemini 2.5 and Claude 4 judges, over the fluent, non-refusing outputs on JailNewsBench test-split fake-news generation instructions (seed instructions built from real news articles of 34 regions in 22 languages, each with a financial, political, social or psychological motive), averaged across regions, under the Context Overload jailbreak (512 tokens of unrelated news articles are prepended to the instruction); lower is better. Source: arxiv.org. 9 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Llama 3.1 8B Instruct | 2.1 | #1018 |
| 2 | DeepSeek R1 Distill Llama 70B | 2.1 | #582 |
| 3 | DeepSeek R1 Distill Llama 8B | 2.1 | #1282 |
| 4 | Llama 3.1 70B Instruct | 2.2 | #548 |
| 5 | GPT-5 | 2.3 | #91 |
| 6 | Claude Sonnet 4 | 2.4 | #194 |
| 7 | Gemini 2.5 Flash | 2.9 | #237 |
Interactive version: theaggregate.ai/benchmark?slug=jailnewsbench-context-overload-harmfulness-score · How It Works · Data refreshed daily, snapshot 2026-10-11.