CaST-Bench (Full Set) - Temporal Grounding: leaderboard

Metric: IM-tIoU (%): mean temporal IoU of predicted evidence segments greedily matched one-to-one to the annotated causal-chain instances, unmatched instances counting zero, all 2,066 six-option questions over 1,015 videos; every model must first output a grounded spatio-temporal causal chain (evidence segments with per-second boxes) and then answer; native video input except the GPT-5 models, which get 1 fps frames; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 13 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash27.63
2GPT-526.61
3Gemini 2.5 Pro21.53
4GPT-5 Mini19.89
5GLM-4.1V-9B (Thinking)16.58
6Qwen 3 VL 4B Instruct11.21
7Qwen 3 VL 8B Instruct10.51
8Qwen 3 VL 8B (Thinking)10.23
9Qwen 2.5 VL 7B Instruct3.72

Interactive version: theaggregate.ai/benchmark?slug=cast-bench-full-set-temporal-grounding · How It Works · Data refreshed daily, snapshot 2026-10-07.