CaST-Bench (Full Set): leaderboard

Metric: Answer accuracy (%) on the multiple-choice questions (chance 16.67), all 2,066 six-option questions over 1,015 videos; every model must first output a grounded spatio-temporal causal chain (evidence segments with per-second boxes) and then answer; native video input except the GPT-5 models, which get 1 fps frames; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 15 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro50.34
2GPT-546.32
3Gemini 2.5 Flash45.6
4Qwen 3 VL 4B Instruct45.3
5Qwen 3 VL 8B Instruct43.13
6Qwen 2.5 VL 7B Instruct41.09
7GLM-4.1V-9B (Thinking)39.55
8Qwen 3 VL 8B (Thinking)39.16
9GPT-5 Mini37.22

Interactive version: theaggregate.ai/benchmark?slug=cast-bench-full-set · How It Works · Data refreshed daily, snapshot 2026-10-07.