ForesightSafety Bench (Industrial Safety, Jailbreak) - Education and Research: leaderboard

Metric: Attack success rate (%) in the education and research sector of ForesightSafetyBench-IndustrialSafety-O: each harmful request is wrapped in the paper's templated role-play jailbreak prompts from JailbreakChat (AIM, DAN and developer-mode templates), and the rate is taken over all attacked prompts; a success is a response that complies with the sector-specific harmful request, judged by gpt-4o-2024-11-20 with dimension-specific rubrics; lower is better. Source: arxiv.org. 18 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.50#138
2Claude Haiku 4.50#271
3Gemini 3 Flash (Preview)0#78
4Qwen 2.5 72B Instruct0.67#436
5Kimi K2.51.33#139
6Seed-1.61.33#257
7Gemini 3 Pro (Preview)2#64
8Kimi K2 09053.33#282
9Seed 1.83.33#136
10Qwen 3 235B A22B 2507 Instruct4#291
11GLM-4.74#185
12Claude 3.5 Haiku4#553
13DeepSeek V3.27.33#198
14Qwen 3 Max (Thinking)7.33#201 (Qwen 3 Max)
15Llama 3.3 70B Instruct12#520

Interactive version: theaggregate.ai/benchmark?slug=foresightsafety-bench-industrial-safety-jailbreak-education-and-research · How It Works · Data refreshed daily, snapshot 2026-10-11.