MT-JailBench - XTeaming: leaderboard

Metric: Attack success rate (%) of the XTeaming multi-turn jailbreak (planned attack strategies with TextGrad prompt refinement) over the 159 HarmBench behaviors used in prior multi-turn jailbreak work, black-box text-only, Qwen-2.5-32B as the attacker model, success judged by the GPT-4o 5-point Score-Judge, at most 5 turns and 20 target interactions with up to 3 retries per turn and no restarts; a lower rate means a more robust target, so lower is better. Source: arxiv.org. 21 models tracked.

Top models

#ModelScore
1Claude Haiku 4.50.63
2Claude Sonnet 4.5 (Thinking)3.14
3Qwen 3.5 Flash (Thinking)23.27
4GPT-532.7
5Qwen 3.5 Plus (Thinking)39.62
6Gemini 3 Pro40.25
7gemma-4-E4B-it46.45
8Llama 3 8B Instruct47.77
9Gemma 4 26B A4B (IT)47.8
10Gemini 3 Flash53.46
11Llama 4 Maverick55.97
12Llama 4 Scout59.75
13Grok 4.1 Fast (Reasoning)62.26
14DeepSeek V3.2 (Thinking)72.96
15Llama 3 70B Instruct74.17

Interactive version: theaggregate.ai/benchmark?slug=mt-jailbench-xteaming · How It Works · Data refreshed daily, snapshot 2026-10-07.