CaST-Bench (Full Set) - Temporal Grounding: leaderboard
Metric: IM-tIoU (%): mean temporal IoU of predicted evidence segments greedily matched one-to-one to the annotated causal-chain instances, unmatched instances counting zero, all 2,066 six-option questions over 1,015 videos; every model must first output a grounded spatio-temporal causal chain (evidence segments with per-second boxes) and then answer; native video input except the GPT-5 models, which get 1 fps frames; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 27.63 |
| 2 | GPT-5 | 26.61 |
| 3 | Gemini 2.5 Pro | 21.53 |
| 4 | GPT-5 Mini | 19.89 |
| 5 | GLM-4.1V-9B (Thinking) | 16.58 |
| 6 | Qwen 3 VL 4B Instruct | 11.21 |
| 7 | Qwen 3 VL 8B Instruct | 10.51 |
| 8 | Qwen 3 VL 8B (Thinking) | 10.23 |
| 9 | Qwen 2.5 VL 7B Instruct | 3.72 |
Interactive version: theaggregate.ai/benchmark?slug=cast-bench-full-set-temporal-grounding · How It Works · Data refreshed daily, snapshot 2026-10-07.