VistaQA - Answer Accuracy: leaderboard

Metric: Overall text answer accuracy (%), on the 1,157 VistaQA samples (six task types, six visual domains, 314 hallucination samples whose queried entity is absent), zero-shot with standardized output-format instructions; the answer is judged correct or not by Qwen 2.5-14B and the evidence masks are matched to the reference masks by Hungarian IoU; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.

Top models

#ModelScore
1SAM3 + Gemini 3 (VistaQA checkpoint unspecified)62.32
2SAM3 + GPT-5.4 (Thinking)61.02
3SAM3 + Qwen3-VL-32B-Instruct57.65
4SAM3 + GPT-5.453.5
5SAM3 + Qwen3-VL-4B-Instruct53.15
6R-Sa2VA-Qwen3VL-4B-RL36.3
7Sa2VA-8B29.04
8UniPixel-7B21.26
9TreeVGR-7B16.08
10LISA-7B7.26

Interactive version: theaggregate.ai/benchmark?slug=vistaqa-answer-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.