VERHallu - Temporal QA: leaderboard

Metric: Accuracy (%; 212 seven-option questions on what happens after an event, distractors from language and vision-language priors; random 14.3; each model's own frame sampling and default settings, answer given as the option number). Source: arxiv.org. Saturation forecast: Around 2029. 18 models tracked.

Top models

#ModelScore
1Gemini 3 Pro55
2Qwen 2.5 VL 7B Instruct50.5
3InternVL3-8B42.9
4Qwen 2.5 VL 32B Instruct34.4
5MiniCPM-V-2.633.9
6Qwen 2.5 VL 72B Instruct28.7

Interactive version: theaggregate.ai/benchmark?slug=verhallu-temporal-qa · How It Works · Data refreshed daily, snapshot 2026-09-26.