VISTA (Video Grounding) - Temporal: leaderboard

Metric: Mean spatio-temporal IoU (m_vIoU, %): per-frame box IoU summed over the frames where predicted and true tubes overlap in time, divided by their temporal union, on the samples the taxonomy labels temporal (entity state transitions over time), over the 11,814 video-query pairs of VISTA (HCSTVG-v1 and v2, VidVRD, VidSTG, MeViS and RVOS videos re-annotated with an interaction taxonomy), zero-shot on sub-sampled frames; the model must localize the queried subject with a box in every frame; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScore
1Qwen3-VL-8B64.3
2InternVL2.5-8B48
3VISTA CogVLM Grounding (checkpoint unspecified)45.7
4Qwen-VL-Chat45.3
5VISTA Sphinx-v2 (checkpoint unspecified)45
6MiniGPT-v244.3
7VISTA MiMo-VL-7B (checkpoint unspecified)43.5
8Shikra-7B32.4
9LLaVA-Grounding-7B31.9
10Ferret-7B23.8

Interactive version: theaggregate.ai/benchmark?slug=vista-video-grounding-temporal · How It Works · Data refreshed daily, snapshot 2026-10-07.