VidPair-Halluc: leaderboard
Metric: Binary-QA weighted accuracy wAcc (%): question-pair and video-pair accuracy (all answers in an adversarial pair correct) weighted by their sample sizes on VidPair-Halluc adversarial video pairs (generated clips with near-identical backgrounds and different foreground semantics). Source: arxiv.org. Saturation forecast: Around 2030. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 49.15 |
| 2 | Qwen 2.5 VL 7B Instruct | 41.66 |
| 3 | Gemini 2.5 Flash | 29.83 |
| 4 | GPT-5 Mini | 29.33 |
| 5 | GPT-4o | 26.97 |
Interactive version: theaggregate.ai/benchmark?slug=vidpair-halluc · How It Works · Data refreshed daily, snapshot 2026-09-29.