VISTA (Video Grounding) - Referral Queries: leaderboard
Metric: Mean spatio-temporal IoU (m_vIoU, %): per-frame box IoU summed over the frames where predicted and true tubes overlap in time, divided by their temporal union, on referral queries (the subject and its attributes only, extracted from the caption by an LLM), over the 11,814 video-query pairs of VISTA (HCSTVG-v1 and v2, VidVRD, VidSTG, MeViS and RVOS videos re-annotated with an interaction taxonomy), zero-shot on sub-sampled frames; the model must localize the queried subject with a box in every frame; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3-VL-8B | 62.85 |
| 2 | VISTA CogVLM Grounding (checkpoint unspecified) | 60.56 |
| 3 | InternVL2.5-8B | 51.11 |
| 4 | VISTA Sphinx-v2 (checkpoint unspecified) | 47.79 |
| 5 | MiniGPT-v2 | 46.62 |
| 6 | Qwen-VL-Chat | 45.56 |
| 7 | VISTA MiMo-VL-7B (checkpoint unspecified) | 43.34 |
| 8 | Shikra-7B | 30.91 |
| 9 | LLaVA-Grounding-7B | 22.51 |
| 10 | Ferret-7B | 17.74 |
Interactive version: theaggregate.ai/benchmark?slug=vista-video-grounding-referral-queries · How It Works · Data refreshed daily, snapshot 2026-10-07.