CCTVBench: leaderboard

Metric: Quadruple accuracy (%): share of question pairs in which the model answers Yes to the true hypothesis on the real video and No to the other three video-question combinations, on CCTVBench's 305 contrastive traffic scenes (a real accident video paired with a world-model-generated counterfactual video) and 1,776 pairs of mutually exclusive yes/no hypothesis questions (7,104 binary instances), single-word Yes/No answers at temperature 0, each question in its own request; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 10 models tracked.

Top models

#ModelScore
1Phi-4 Multimodal Instruct2.92

Interactive version: theaggregate.ai/benchmark?slug=cctvbench · How It Works · Data refreshed daily, snapshot 2026-10-07.