UniSAFE - Multimodal Understanding ASR: leaderboard

Metric: Attack success rate (%): share of UniSAFE's unsafe-intent prompts (benign-looking inputs built to elicit harmful content across safety categories) whose output an ensemble of three MLLM judges (Gemini-2.5 Pro, GPT-5-nano, Qwen2.5-VL-72B) rates harmful, on the multimodal understanding with text output task; lower is better. Source: arxiv.org. 13 models tracked.

Top models

#ModelScoreOverall rank
1GPT-512.7#91
2Gemini 2.5 Pro36#145
3Qwen 2.5 VL 7B Instruct43.6#643

Interactive version: theaggregate.ai/benchmark?slug=unisafe-multimodal-understanding-asr · How It Works · Data refreshed daily, snapshot 2026-10-11.