VISTA (Video Grounding) - Spatial: leaderboard
Metric: Mean spatio-temporal IoU (m_vIoU, %): per-frame box IoU summed over the frames where predicted and true tubes overlap in time, divided by their temporal union, on the samples the taxonomy labels spatial (positional configurations among entities), over the 11,814 video-query pairs of VISTA (HCSTVG-v1 and v2, VidVRD, VidSTG, MeViS and RVOS videos re-annotated with an interaction taxonomy), zero-shot on sub-sampled frames; the model must localize the queried subject with a box in every frame; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3-VL-8B | 64.8 |
| 2 | VISTA CogVLM Grounding (checkpoint unspecified) | 57.5 |
| 3 | InternVL2.5-8B | 46.3 |
| 4 | Qwen-VL-Chat | 45.7 |
| 5 | MiniGPT-v2 | 43.1 |
| 6 | VISTA Sphinx-v2 (checkpoint unspecified) | 42.6 |
| 7 | VISTA MiMo-VL-7B (checkpoint unspecified) | 36.9 |
| 8 | Shikra-7B | 29.9 |
| 9 | LLaVA-Grounding-7B | 28.1 |
| 10 | Ferret-7B | 20.9 |
Interactive version: theaggregate.ai/benchmark?slug=vista-video-grounding-spatial · How It Works · Data refreshed daily, snapshot 2026-10-07.