TukaBench - English: leaderboard

Metric: Attack success rate (%; share of harmful prompts answered Jailbroken) on the English versions of the TukaBench jailbreak prompts (JailbreakBench-derived Afri-JBB-Harm and Afri-JBB-Culture plus African-authored AfriJail-Mono), direct prompting, greedy decoding, GPT-4.1 judge labelling each response Jailbroken, Refused or Deflected; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)2.1
2GPT-OSS-120B4.7
3Qwen 3.5 27B5.4
4GPT-5.25.8
5Claude Sonnet 59.4
6Claude Opus 4.810.9
7DeepSeek V3.211.4
8Claude Haiku 4.513.9
9Gemma 4 31B15.4
10GPT-4o16.5
11Llama 4 Maverick17.1
12Grok 4.317.9
13Grok 423.1
14Gemma 3 27B26.1
15Grok 332.8

Interactive version: theaggregate.ai/benchmark?slug=tukabench-english · How It Works · Data refreshed daily, snapshot 2026-09-29.