KorSET (Direct Request) - Misinformation Harms: leaderboard
Metric: Attack success rate (%) of KorSET's Misinformation Harms prompts sent as direct requests (no adversarial optimization) to the target model; a GPT-4.1 judge with the authors' safety rubric decides whether each response is harmful; lower is better. Source: arxiv.org. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemma 3 12B (IT) | 12.27 | #655 |
| 2 | EXAONE 3.5 7.8B Instruct | 13.85 | #850 |
| 3 | Gemma 2 9B (IT) | 20.92 | #774 |
| 4 | Qwen 2.5 7B Instruct | 21.16 | #846 |
| 5 | Llama 3.1 8B Instruct | 48.78 | #1018 |
Interactive version: theaggregate.ai/benchmark?slug=korset-direct-request-misinformation-harms · How It Works · Data refreshed daily, snapshot 2026-10-11.