CCTVBench (Video-Level Batch): leaderboard

Metric: Quadruple accuracy (%): share of question pairs in which the model answers Yes to the true hypothesis on the real video and No to the other three video-question combinations, on CCTVBench's 305 contrastive traffic scenes (a real accident video paired with a world-model-generated counterfactual video) and 1,776 pairs of mutually exclusive yes/no hypothesis questions (7,104 binary instances), single-word Yes/No answers at temperature 0, all of a video's questions in one video-level batch request, each answer parsed separately (the paper's protocol for proprietary APIs); higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 4 models tracked.

Top models

#ModelScore
1GPT-5 Mini25.85
2Gemini 2.5 Flash20.87
3Gemini 2.5 Flash Lite18.28
4GPT-5 Nano12.6

Interactive version: theaggregate.ai/benchmark?slug=cctvbench-video-level-batch · How It Works · Data refreshed daily, snapshot 2026-10-07.