SaLAD: leaderboard

Metric: Accuracy rate (%; GPT-4o judge; all 2,013 image-text queries: an unsafe query is correct when the answer names and explains the risk, a safe query when it is answered without refusal or needless warnings). Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet77.05
2Gemini 2.5 Flash75.96
3GPT-4o75.36
4Qwen 2.5 VL 7B62.89
5Llama 3.2 11B Vision61.85
6MiniCPM-V-2.659.12

Interactive version: theaggregate.ai/benchmark?slug=salad · How It Works · Data refreshed daily, snapshot 2026-09-26.