VidPair-Halluc: leaderboard

Metric: Binary-QA weighted accuracy wAcc (%): question-pair and video-pair accuracy (all answers in an adversarial pair correct) weighted by their sample sizes on VidPair-Halluc adversarial video pairs (generated clips with near-identical backgrounds and different foreground semantics). Source: arxiv.org. Saturation forecast: Around 2030. 15 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro49.15
2Qwen 2.5 VL 7B Instruct41.66
3Gemini 2.5 Flash29.83
4GPT-5 Mini29.33
5GPT-4o26.97

Interactive version: theaggregate.ai/benchmark?slug=vidpair-halluc · How It Works · Data refreshed daily, snapshot 2026-09-29.