ProactiveBench: leaderboard

Metric: Trajectory accuracy (%) averaged over the seven scenarios (ROD, VSOD, MVP-N, ImageNet-C, QuickDraw, ChangeIt, COCO): multi-turn multiple choice in which the model may name a category, abstain, or pick a proactive suggestion that yields a new frame, and a trajectory counts only when it ends in the correct category; samples a first-turn guess solves for at least 25% of the evaluated models are filtered out; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 22 models tracked.

Top models

#ModelScoreOverall rank
1O4 Mini34#172
2GPT-4.130.9#240
3GPT-5.224.8#105
4InternVL3-38B23#395
5Phi-4 Multimodal Instruct19.4#896
6InternVL3-78B15.6#345
7InternVL3-8B12.7#606
8Qwen 2.5 VL 32B Instruct10.8#443
9Qwen 2.5 VL 7B Instruct9.9#643
10Qwen 2.5 VL 72B Instruct7.5#364

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=proactivebench · How It Works · Data refreshed daily, snapshot 2026-10-11.