ProactiveBench - VSOD: leaderboard

Metric: Trajectory accuracy (%) on the temporally occluded video frames of VSOD: multi-turn multiple choice in which the model may name a category, abstain, or pick a proactive suggestion that yields a new frame, and a trajectory counts only when it ends in the correct category; samples a first-turn guess solves for at least 25% of the evaluated models are filtered out; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 22 models tracked.

Top models

#ModelScoreOverall rank
1InternVL3-38B31#395
2InternVL3-78B16.7#345
3O4 Mini16.7#172
4Phi-4 Multimodal Instruct16.7#896
5InternVL3-8B11.9#606
6Qwen 2.5 VL 32B Instruct4.8#443
7Qwen 2.5 VL 72B Instruct2.4#364
8GPT-5.20#105
9GPT-4.10#240
10Qwen 2.5 VL 7B Instruct0#643

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=proactivebench-vsod · How It Works · Data refreshed daily, snapshot 2026-10-11.