CaST-Bench (Full Set) - Open-Ended: leaderboard
Metric: Open-ended justification score (0-10): an LLM judge rates the free-form answer and causal chain against the ground truth for answer correctness, logical consistency, evidence coverage and overall justification, all 2,066 six-option questions over 1,015 videos; every model must first output a grounded spatio-temporal causal chain (evidence segments with per-second boxes) and then answer; native video input except the GPT-5 models, which get 1 fps frames; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 4.3 |
| 2 | GPT-5 Mini | 3.98 |
| 3 | Gemini 2.5 Pro | 3.85 |
| 4 | Gemini 2.5 Flash | 3.83 |
| 5 | GLM-4.1V-9B (Thinking) | 2.66 |
| 6 | Qwen 3 VL 4B Instruct | 2.56 |
| 7 | Qwen 3 VL 8B (Thinking) | 2.15 |
| 8 | Qwen 2.5 VL 7B Instruct | 1.94 |
| 9 | Qwen 3 VL 8B Instruct | 1.93 |
Interactive version: theaggregate.ai/benchmark?slug=cast-bench-full-set-open-ended · How It Works · Data refreshed daily, snapshot 2026-10-07.