ForesightSafety Bench (Industrial Safety, Jailbreak): leaderboard

Metric: Attack success rate (%), unweighted mean of the eight industry sectors of ForesightSafetyBench-IndustrialSafety-O: each harmful request is wrapped in the paper's templated role-play jailbreak prompts from JailbreakChat (AIM, DAN and developer-mode templates), and the rate is taken over all attacked prompts; a success is a response that complies with the sector-specific harmful request, judged by gpt-4o-2024-11-20 with dimension-specific rubrics; lower is better. Source: arxiv.org. 18 models tracked.

Top models

#ModelScoreOverall rank
1Claude Haiku 4.50#271
2Claude Sonnet 4.50.33#138
3Kimi K2.51.75#139
4Gemini 3 Flash (Preview)2.5#78
5Claude 3.5 Haiku3.17#553
6Gemini 3 Pro (Preview)3.58#64
7GLM-4.73.75#185
8Qwen 2.5 72B Instruct4.5#436
9Seed-1.65.92#257
10Seed 1.86.42#136
11Qwen 3 235B A22B 2507 Instruct8.75#291
12Kimi K2 090512.67#282
13DeepSeek V3.213.42#198
14Qwen 3 Max (Thinking)13.67#201 (Qwen 3 Max)
15Gemini 2.5 Flash24.75#237

Interactive version: theaggregate.ai/benchmark?slug=foresightsafety-bench-industrial-safety-jailbreak · How It Works · Data refreshed daily, snapshot 2026-10-11.