CaST-Bench (Full Set): leaderboard
Metric: Answer accuracy (%) on the multiple-choice questions (chance 16.67), all 2,066 six-option questions over 1,015 videos; every model must first output a grounded spatio-temporal causal chain (evidence segments with per-second boxes) and then answer; native video input except the GPT-5 models, which get 1 fps frames; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 50.34 |
| 2 | GPT-5 | 46.32 |
| 3 | Gemini 2.5 Flash | 45.6 |
| 4 | Qwen 3 VL 4B Instruct | 45.3 |
| 5 | Qwen 3 VL 8B Instruct | 43.13 |
| 6 | Qwen 2.5 VL 7B Instruct | 41.09 |
| 7 | GLM-4.1V-9B (Thinking) | 39.55 |
| 8 | Qwen 3 VL 8B (Thinking) | 39.16 |
| 9 | GPT-5 Mini | 37.22 |
Interactive version: theaggregate.ai/benchmark?slug=cast-bench-full-set · How It Works · Data refreshed daily, snapshot 2026-10-07.