MIR-SafetyBench - Temporal Continuity: leaderboard

Metric: Attack Success Rate (%; lower is safer). Source: arxiv.org. Saturation forecast: Around December 2026. 19 models tracked.

Top models

#ModelScore
1GPT-5.126.18
2Gemini 3 Pro (Preview)53.63
3Gemini 2.5 Pro61.51
4GPT-4o Mini65.62
5QVQ-72B-Preview72.24
6GPT-4o74.76
7Gemini 2.5 Flash76.34
8InternVL3-38B79.5
9InternVL3-8B79.81
10InternVL3-78B83.91
11Qwen 2.5 VL 32B Instruct85.17
12GLM-4.1V-9B (Thinking)85.49

Interactive version: theaggregate.ai/benchmark?slug=mir-safetybench-temporal-continuity · How It Works · Data refreshed daily, snapshot 2026-09-25.