FCMBench-Video (Chinese) - Evidence-Grounded Selection: leaderboard
Metric: Evidence-Grounded Selection: exact-match accuracy (%) on single-choice questions answered from the right document in a multi-document video, on FCMBench-Video's zh-CN subset (405 composed videos of 20 Chinese financial-credit document types, 5,960 instances); zero-shot with structured outputs; each model's native video serving path (2 FPS sampling for most open models, 0.5 FPS for Ovis2.5, raw video upload for the API models); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 73.04 |
| 2 | Qwen 3.5 27B | 67.32 |
| 3 | Qwen 3 VL 32B Instruct | 44.11 |
| 4 | Qwen 3 VL 8B Instruct | 35 |
| 5 | Qwen3 Omni 30B A3B Instruct | 33.75 |
| 6 | InternVL3-8B | 15.18 |
Interactive version: theaggregate.ai/benchmark?slug=fcmbench-video-chinese-evidence-grounded-selection · How It Works · Data refreshed daily, snapshot 2026-10-07.