RoboProcessBench - Next Primitive Prediction: leaderboard

Metric: Accuracy (%) on 200 questions asking which low-level primitive should happen next (4 options, chance 25.0); zero-shot multiple-choice VQA on held-out robot manipulation recordings from GM-100, RH20T, REASSEMBLE and AIST-Bimanual (ProcessData-Eval), temperature 0.01; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 14 models tracked.

Top models

#ModelScore
1InternVL3-38B63.5
2GLM-4.6V60
3Claude Haiku 4.556
4InternVL3.5-8B53.5
5Qwen 3 VL 32B47
6InternVL3-8B46.5
7GPT-5.4 Mini45.5
8Claude Sonnet 4.644.8
9GPT-4o44.5
10Qwen 2.5 VL 7B Instruct33.1

Interactive version: theaggregate.ai/benchmark?slug=roboprocessbench-next-primitive-prediction · How It Works · Data refreshed daily, snapshot 2026-09-29.