SpatialTrust: leaderboard
Metric: Overall score (%; mean over the five question types of sensitive-factor detection, direct and indirect factor identification and explanation; over SpatialTrust's 577 secure-authentication images (five questions each); IoU threshold 0.5 for grounded evidence). Source: arxiv.org. Saturation forecast: Around 2031. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 VL 30B A3B Instruct | 36.78 |
| 2 | Gemini 2.5 Flash | 33.81 |
| 3 | Qwen 3 VL 8B Instruct | 28.91 |
| 4 | GPT-5.4 | 25.94 |
| 5 | GLM-4.1V-9B (Thinking) | 21.62 |
| 6 | Qwen 2 VL 7B Instruct | 15.89 |
Interactive version: theaggregate.ai/benchmark?slug=spatialtrust · How It Works · Data refreshed daily, snapshot 2026-09-26.