EG-VQA - Causal: leaderboard

Metric: Strict accuracy (%) on causal questions (why an event happens or what it causes), on the 2,889-question EG-VQA test split (open-ended questions over ActivityNet Captions, HiREST and YouCook2 videos, 64 uniformly sampled frames), answers graded by a Gemini-2.5-Pro judge against the reference answer; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScore
1GPT-4o37.2
2Gemini 2.5 Flash33.93
3Qwen 2.5 VL 7B Instruct22.36
4InternVL3-8B18.11

Interactive version: theaggregate.ai/benchmark?slug=eg-vqa-causal · How It Works · Data refreshed daily, snapshot 2026-09-29.