MT-JailBench - XTeaming: leaderboard
Metric: Attack success rate (%) of the XTeaming multi-turn jailbreak (planned attack strategies with TextGrad prompt refinement) over the 159 HarmBench behaviors used in prior multi-turn jailbreak work, black-box text-only, Qwen-2.5-32B as the attacker model, success judged by the GPT-4o 5-point Score-Judge, at most 5 turns and 20 target interactions with up to 3 retries per turn and no restarts; a lower rate means a more robust target, so lower is better. Source: arxiv.org. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Haiku 4.5 | 0.63 |
| 2 | Claude Sonnet 4.5 (Thinking) | 3.14 |
| 3 | Qwen 3.5 Flash (Thinking) | 23.27 |
| 4 | GPT-5 | 32.7 |
| 5 | Qwen 3.5 Plus (Thinking) | 39.62 |
| 6 | Gemini 3 Pro | 40.25 |
| 7 | gemma-4-E4B-it | 46.45 |
| 8 | Llama 3 8B Instruct | 47.77 |
| 9 | Gemma 4 26B A4B (IT) | 47.8 |
| 10 | Gemini 3 Flash | 53.46 |
| 11 | Llama 4 Maverick | 55.97 |
| 12 | Llama 4 Scout | 59.75 |
| 13 | Grok 4.1 Fast (Reasoning) | 62.26 |
| 14 | DeepSeek V3.2 (Thinking) | 72.96 |
| 15 | Llama 3 70B Instruct | 74.17 |
Interactive version: theaggregate.ai/benchmark?slug=mt-jailbench-xteaming · How It Works · Data refreshed daily, snapshot 2026-10-07.