CaST-Bench (Full Set) - Open-Ended: leaderboard

Metric: Open-ended justification score (0-10): an LLM judge rates the free-form answer and causal chain against the ground truth for answer correctness, logical consistency, evidence coverage and overall justification, all 2,066 six-option questions over 1,015 videos; every model must first output a grounded spatio-temporal causal chain (evidence segments with per-second boxes) and then answer; native video input except the GPT-5 models, which get 1 fps frames; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 15 models tracked.

Top models

#ModelScore
1GPT-54.3
2GPT-5 Mini3.98
3Gemini 2.5 Pro3.85
4Gemini 2.5 Flash3.83
5GLM-4.1V-9B (Thinking)2.66
6Qwen 3 VL 4B Instruct2.56
7Qwen 3 VL 8B (Thinking)2.15
8Qwen 2.5 VL 7B Instruct1.94
9Qwen 3 VL 8B Instruct1.93

Interactive version: theaggregate.ai/benchmark?slug=cast-bench-full-set-open-ended · How It Works · Data refreshed daily, snapshot 2026-10-07.