KorSET (Direct Request) - Misinformation Harms: leaderboard

Metric: Attack success rate (%) of KorSET's Misinformation Harms prompts sent as direct requests (no adversarial optimization) to the target model; a GPT-4.1 judge with the authors' safety rubric decides whether each response is harmful; lower is better. Source: arxiv.org. 5 models tracked.

Top models

#ModelScoreOverall rank
1Gemma 3 12B (IT)12.27#655
2EXAONE 3.5 7.8B Instruct13.85#850
3Gemma 2 9B (IT)20.92#774
4Qwen 2.5 7B Instruct21.16#846
5Llama 3.1 8B Instruct48.78#1018

Interactive version: theaggregate.ai/benchmark?slug=korset-direct-request-misinformation-harms · How It Works · Data refreshed daily, snapshot 2026-10-11.