FCMBench-Video (English) - Evidence-Grounded Selection: leaderboard
Metric: Evidence-Grounded Selection: exact-match accuracy (%) on single-choice questions answered from the right document in a multi-document video, on FCMBench-Video's en-US subset (795 composed videos of 8 English document types, 5,362 instances); zero-shot with structured outputs; each model's native video serving path (2 FPS sampling for most open models, 0.5 FPS for Ovis2.5, raw video upload for the API models); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 27B | 86.85 |
| 2 | Gemini 3 Pro (Preview) | 76.99 |
| 3 | Qwen 3 VL 32B Instruct | 41.9 |
| 4 | Qwen3 Omni 30B A3B Instruct | 38.53 |
| 5 | Qwen 3 VL 8B Instruct | 33.03 |
| 6 | InternVL3-8B | 32.72 |
Interactive version: theaggregate.ai/benchmark?slug=fcmbench-video-english-evidence-grounded-selection · How It Works · Data refreshed daily, snapshot 2026-10-07.