RedVox: leaderboard

Metric: Unsafe response rate (%; share of responses a GPT-5.5 judge labels unsafe (endorsing, supporting or complying with the harmful or stereotyped request, whether or not the request was understood); all five languages and request types of the 3,414 released entries). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Gemini 3.1 Flash Lite2.6
2Gemini 3.1 Pro (Preview)3.1
3Qwen3 Omni 30B A3B Instruct3.4
4Qwen2-Audio-7B-Instruct10.9
5Phi-4 Multimodal Instruct16.1
6Voxtral-Small-24B-250721.9

Interactive version: theaggregate.ai/benchmark?slug=redvox · How It Works · Data refreshed daily, snapshot 2026-09-26.