VistaQA - Hallucination Samples: leaderboard

Metric: Grove score (10-100): per sample the geometric mean of the answer score and the mask score, each floored at 0.1, averaged over samples and times 100, on the 314 hallucination samples whose queried entity is absent (the correct mask is none), on the 1,157 VistaQA samples (six task types, six visual domains, 314 hallucination samples whose queried entity is absent), zero-shot with standardized output-format instructions; the answer is judged correct or not by Qwen 2.5-14B and the evidence masks are matched to the reference masks by Hungarian IoU; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.

Top models

#ModelScore
1SAM3 + Qwen3-VL-4B-Instruct60.72
2SAM3 + Qwen3-VL-32B-Instruct56.41
3SAM3 + Gemini 3 (VistaQA checkpoint unspecified)56.14
4SAM3 + GPT-5.4 (Thinking)47.89
5SAM3 + GPT-5.436.72
6SESAME-7B30.04
7TreeVGR-7B22.66
8R-Sa2VA-Qwen3VL-4B-RL20.79
9LISA-7B19.39
10UGround-7B14.03

Interactive version: theaggregate.ai/benchmark?slug=vistaqa-hallucination-samples · How It Works · Data refreshed daily, snapshot 2026-10-07.