MMJailBench - System-Role Framing: leaderboard

Metric: Attack success rate (%; share of responses a GPT-5 judge scores 4 or 5 on a 1-5 harmfulness rubric; system-role simulation framing, all 272 intents, visual contexts and carriers). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-50
2Claude Sonnet 4.50
3InternVL3-8B0.07
4Qwen 2.5 VL 7B Instruct0.66
5Gemini 3 Flash3.2
6Grok 4 Fast3.53
7Step3 VL 10B15.51
8Seed 2.0 Lite16.1
9Gemma 3 12B18.9
10Qwen 3 VL 8B22.32
11Ministral 3 8B30.7
12Kimi K2.540.33
13GLM-4.6V79.56

Interactive version: theaggregate.ai/benchmark?slug=mmjailbench-system-role-framing · How It Works · Data refreshed daily, snapshot 2026-09-26.