EG-VQA - Counterfactual: leaderboard
Metric: Strict accuracy (%) on counterfactual questions (outcomes of a hypothetical change), on the 2,889-question EG-VQA test split (open-ended questions over ActivityNet Captions, HiREST and YouCook2 videos, 64 uniformly sampled frames), answers graded by a Gemini-2.5-Pro judge against the reference answer; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 35.1 |
| 2 | Gemini 2.5 Flash | 28.89 |
| 3 | Qwen 2.5 VL 7B Instruct | 16.8 |
| 4 | InternVL3-8B | 13.64 |
Interactive version: theaggregate.ai/benchmark?slug=eg-vqa-counterfactual · How It Works · Data refreshed daily, snapshot 2026-09-29.