FCMBench-Video (English) - Evidence-Grounded Selection: leaderboard

Metric: Evidence-Grounded Selection: exact-match accuracy (%) on single-choice questions answered from the right document in a multi-document video, on FCMBench-Video's en-US subset (795 composed videos of 8 English document types, 5,362 instances); zero-shot with structured outputs; each model's native video serving path (2 FPS sampling for most open models, 0.5 FPS for Ovis2.5, raw video upload for the API models); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B86.85
2Gemini 3 Pro (Preview)76.99
3Qwen 3 VL 32B Instruct41.9
4Qwen3 Omni 30B A3B Instruct38.53
5Qwen 3 VL 8B Instruct33.03
6InternVL3-8B32.72

Interactive version: theaggregate.ai/benchmark?slug=fcmbench-video-english-evidence-grounded-selection · How It Works · Data refreshed daily, snapshot 2026-10-07.