VistaQA - Evidence Mask mIoU: leaderboard

Metric: Overall evidence mask score (%): mean per-sample mask score (Hungarian-matched IoU, 1 for correctly predicting no mask on a hallucination sample), on the 1,157 VistaQA samples (six task types, six visual domains, 314 hallucination samples whose queried entity is absent), zero-shot with standardized output-format instructions; the answer is judged correct or not by Qwen 2.5-14B and the evidence masks are matched to the reference masks by Hungarian IoU; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.

Top models

#ModelScore
1SAM3 + GPT-5.4 (Thinking)41.98
2SAM3 + Qwen3-VL-4B-Instruct39.91
3SAM3 + GPT-5.439.26
4SAM3 + Qwen3-VL-32B-Instruct36.02
5SAM3 + Gemini 3 (VistaQA checkpoint unspecified)34.99
6R-Sa2VA-Qwen3VL-4B-RL24.25
7SESAME-7B23.67
8Sa2VA-8B21.01
9UniPixel-7B20.43
10TreeVGR-7B19.66

Interactive version: theaggregate.ai/benchmark?slug=vistaqa-evidence-mask-miou · How It Works · Data refreshed daily, snapshot 2026-10-07.