MME-Safety - Harmful Rate: leaderboard

Metric: Harmful rate (%; share of final answers to 3,801 manually verified image-text pairs with harmful intent across nine risk categories and nine image/text stealth combinations, zero-shot; GPT-4o judge whose response contains harmful content; lower is better). Source: arxiv.org. Saturation forecast: Around December 2026. 17 models tracked.

Top models

#ModelScore
1Qwen 3 VL 235B A22B Instruct11.31
2Claude Sonnet 4 (Thinking)11.87
3Claude 3.7 Sonnet12.8
4Claude 3.7 Sonnet (Thinking)13.96
5O4 Mini15.32
6Qwen 3 VL 235B A22B (Thinking)15.34
7Qwen 3 VL 32B (Thinking)17.5
8Qwen 3 VL 8B (Thinking)19.83
9Gemini 3 Pro34.42

Interactive version: theaggregate.ai/benchmark?slug=mme-safety-harmful-rate · How It Works · Data refreshed daily, snapshot 2026-09-26.