RoboProcessBench - Outcome Prediction: leaderboard

Metric: Accuracy (%) on 793 questions asking whether the ongoing attempt will eventually succeed (2 options, chance 50.0); zero-shot multiple-choice VQA on held-out robot manipulation recordings from GM-100, RH20T, REASSEMBLE and AIST-Bimanual (ProcessData-Eval), temperature 0.01; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 14 models tracked.

Top models

#ModelScore
1InternVL3-8B61.9
2InternVL3.5-8B61.9
3GPT-5.4 Mini56.1
4InternVL3-38B55.9
5Qwen 2.5 VL 7B Instruct54.6
6GLM-4.6V53
7Qwen 3 VL 32B50.6
8Claude Sonnet 4.649.7
9Claude Haiku 4.546.2
10GPT-4o41.1

Interactive version: theaggregate.ai/benchmark?slug=roboprocessbench-outcome-prediction · How It Works · Data refreshed daily, snapshot 2026-09-29.