RuleSafe-VL - Decision Relation Identification: leaderboard
Metric: Macro-F1 (%; typed relations among a policy's atomic rules from the policy text, mean of three policy families). Source: arxiv.org. Saturation forecast: Around March 2027. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 64.8 |
| 2 | Claude Opus 4.6 | 63.4 |
| 3 | GLM-5.1 | 55.8 |
| 4 | MiniMax-M2.7 | 50.7 |
| 5 | Qwen 3 VL 8B | 37.1 |
| 6 | Qwen 3 VL 4B | 35.2 |
| 7 | Qwen 2.5 VL 7B Instruct | 25.6 |
Interactive version: theaggregate.ai/benchmark?slug=rulesafe-vl-decision-relation-identification · How It Works · Data refreshed daily, snapshot 2026-09-25.