MMJailBench - Violence & Physical Harm: leaderboard

Metric: Attack success rate (%; share of responses a GPT-5 judge scores 4 or 5 on a 1-5 harmfulness rubric; Violence & Physical Harm domain, 48 harmful intents across all framings, visual contexts and carriers). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-50.07
2Claude Sonnet 4.52.08
3Qwen 3 VL 8B11.15
4Gemini 3 Flash11.94
5Kimi K2.523.3
6Grok 4 Fast26.91
7Seed 2.0 Lite32.99
8Step3 VL 10B44.24
9Ministral 3 8B46.98
10InternVL3-8B50.8
11Gemma 3 12B60.66
12Qwen 2.5 VL 7B Instruct66.91
13GLM-4.6V73.68

Interactive version: theaggregate.ai/benchmark?slug=mmjailbench-violence-physical-harm · How It Works · Data refreshed daily, snapshot 2026-09-26.