RuleSafe-VL - Atomic Rule Identification: leaderboard

Metric: Macro-F1 (%; whether each candidate atomic policy rule is satisfied by an image-text case, mean of three policy families). Source: arxiv.org. Saturation forecast: Around 2028. 9 models tracked.

Top models

#ModelScore
1Claude Opus 4.669.3
2Gemini 3.1 Pro (Preview)68.8
3Qwen 3 VL 8B63
4Qwen 3 VL 4B59.8
5GPT-5.456.8
6Qwen 2.5 VL 7B Instruct47

Interactive version: theaggregate.ai/benchmark?slug=rulesafe-vl-atomic-rule-identification · How It Works · Data refreshed daily, snapshot 2026-09-25.