MIR-SafetyBench - Logical Analogy: leaderboard

Metric: Attack Success Rate (%; lower is safer). Source: arxiv.org. Saturation forecast: Estimated already saturated. 19 models tracked.

Top models

#ModelScore
1GPT-5.15.03
2Gemini 3 Pro (Preview)26.1
3Gemini 2.5 Pro27.36
4Gemini 2.5 Flash42.77
5GPT-4o Mini55.35
6GPT-4o58.49
7InternVL3-78B67.3
8QVQ-72B-Preview67.92
9InternVL3-8B78.62
10InternVL3-38B81.13
11Qwen 2.5 VL 32B Instruct81.76
12GLM-4.1V-9B (Thinking)93.4

Interactive version: theaggregate.ai/benchmark?slug=mir-safetybench-logical-analogy · How It Works · Data refreshed daily, snapshot 2026-09-25.