MT-JailBench - Crescendo: leaderboard
Metric: Attack success rate (%) of the Crescendo multi-turn jailbreak (fixed-template escalating prompts, retry on refusal) over the 159 HarmBench behaviors used in prior multi-turn jailbreak work, black-box text-only, Qwen-2.5-32B as the attacker model, success judged by the GPT-4o 5-point Score-Judge, at most 5 turns and 20 target interactions with up to 3 retries per turn and no restarts; a lower rate means a more robust target, so lower is better. Source: arxiv.org. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Haiku 4.5 | 21.38 |
| 2 | GPT-5 | 25.16 |
| 3 | Qwen 3.5 Flash (Thinking) | 28.3 |
| 4 | Claude Sonnet 4.5 (Thinking) | 30.82 |
| 5 | Qwen 3.5 Plus (Thinking) | 36.48 |
| 6 | gemma-4-E4B-it | 48.43 |
| 7 | Grok 4.1 Fast (Reasoning) | 49.69 |
| 8 | Llama 3 8B Instruct | 57.86 |
| 9 | Gemma 4 26B A4B (IT) | 58.49 |
| 10 | Gemini 3 Pro | 62.89 |
| 11 | Llama 4 Maverick | 63.52 |
| 12 | Llama 4 Scout | 64.78 |
| 13 | DeepSeek V3.2 (Thinking) | 64.78 |
| 14 | Llama 3 70B Instruct | 68.55 |
| 15 | Gemini 3 Flash | 72.96 |
Interactive version: theaggregate.ai/benchmark?slug=mt-jailbench-crescendo · How It Works · Data refreshed daily, snapshot 2026-10-07.