MT-JailBench - Crescendo: leaderboard

Metric: Attack success rate (%) of the Crescendo multi-turn jailbreak (fixed-template escalating prompts, retry on refusal) over the 159 HarmBench behaviors used in prior multi-turn jailbreak work, black-box text-only, Qwen-2.5-32B as the attacker model, success judged by the GPT-4o 5-point Score-Judge, at most 5 turns and 20 target interactions with up to 3 retries per turn and no restarts; a lower rate means a more robust target, so lower is better. Source: arxiv.org. 21 models tracked.

Top models

#ModelScore
1Claude Haiku 4.521.38
2GPT-525.16
3Qwen 3.5 Flash (Thinking)28.3
4Claude Sonnet 4.5 (Thinking)30.82
5Qwen 3.5 Plus (Thinking)36.48
6gemma-4-E4B-it48.43
7Grok 4.1 Fast (Reasoning)49.69
8Llama 3 8B Instruct57.86
9Gemma 4 26B A4B (IT)58.49
10Gemini 3 Pro62.89
11Llama 4 Maverick63.52
12Llama 4 Scout64.78
13DeepSeek V3.2 (Thinking)64.78
14Llama 3 70B Instruct68.55
15Gemini 3 Flash72.96

Interactive version: theaggregate.ai/benchmark?slug=mt-jailbench-crescendo · How It Works · Data refreshed daily, snapshot 2026-10-07.