UniSAFE - Text-to-Text ASR: leaderboard
Metric: Attack success rate (%): share of UniSAFE's unsafe-intent prompts (benign-looking inputs built to elicit harmful content across safety categories) whose output an ensemble of three MLLM judges (Gemini-2.5 Pro, GPT-5-nano, Qwen2.5-VL-72B) rates harmful, on the text-to-text generation task; lower is better. Source: arxiv.org. 13 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 | 6.2 | #91 |
| 2 | Gemini 2.5 Pro | 36.4 | #145 |
| 3 | Qwen 2.5 VL 7B Instruct | 48.7 | #643 |
Interactive version: theaggregate.ai/benchmark?slug=unisafe-text-to-text-asr · How It Works · Data refreshed daily, snapshot 2026-10-11.