EG-VQA - Evidence Grounding: leaderboard
Metric: Evidence-grounded F1 (%, EG-F1 at temporal IoU 0.3 and semantic similarity 0.5): F1 of the predicted evidence segments matched one to one to the reference evidence when both the temporal IoU and the description similarity pass the thresholds, each match weighted by IoU times similarity, on the 2,889-question EG-VQA test split (open-ended questions over ActivityNet Captions, HiREST and YouCook2 videos, 64 uniformly sampled frames); higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 12.64 |
| 2 | GPT-4o | 8.49 |
| 3 | InternVL3-8B | 8.18 |
| 4 | Qwen 2.5 VL 7B Instruct | 5.11 |
Interactive version: theaggregate.ai/benchmark?slug=eg-vqa-evidence-grounding · How It Works · Data refreshed daily, snapshot 2026-09-29.