MMJailBench - Economic Harm: leaderboard

Metric: Attack success rate (%; share of responses a GPT-5 judge scores 4 or 5 on a 1-5 harmfulness rubric; Economic Harm domain, 30 harmful intents across all framings, visual contexts and carriers). Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-54.11
2Claude Sonnet 4.515.11
3Qwen 3 VL 8B30.28
4Gemini 3 Flash38.94
5Kimi K2.545.11
6Grok 4 Fast52.5
7InternVL3-8B61.72
8Seed 2.0 Lite61.94
9Ministral 3 8B63.78
10Step3 VL 10B68.22
11Qwen 2.5 VL 7B Instruct70.83
12Gemma 3 12B80.44
13GLM-4.6V88.11

Interactive version: theaggregate.ai/benchmark?slug=mmjailbench-economic-harm · How It Works · Data refreshed daily, snapshot 2026-09-26.