MME-Safety: leaderboard

Metric: Overall safety score (%; mean of relevance rate, 1 - harmful rate x normalised average harm severity, and refusal rate, all on final answers to 3,801 manually verified image-text pairs with harmful intent across nine risk categories and nine image/text stealth combinations, zero-shot; GPT-4o judge). Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.

Top models

#ModelScore
1Claude Sonnet 4 (Thinking)86.24
2Qwen 3 VL 235B A22B Instruct85.11
3O4 Mini84.78
4Claude 3.7 Sonnet84.42
5Claude 3.7 Sonnet (Thinking)82.17
6Qwen 3 VL 235B A22B (Thinking)81.16
7Qwen 3 VL 32B (Thinking)80.49
8Qwen 3 VL 8B (Thinking)77.33
9Gemini 3 Pro75.55

Interactive version: theaggregate.ai/benchmark?slug=mme-safety · How It Works · Data refreshed daily, snapshot 2026-09-26.