RoboProcessBench - Current Primitive Recognition: leaderboard

Metric: Accuracy (%) on 359 questions asking which low-level primitive is being executed, from a short clip (4 options, chance 25.0); zero-shot multiple-choice VQA on held-out robot manipulation recordings from GM-100, RH20T, REASSEMBLE and AIST-Bimanual (ProcessData-Eval), temperature 0.01; higher is better. Source: arxiv.org. Saturation forecast: Around 2035. 14 models tracked.

Top models

#ModelScore
1InternVL3.5-8B36.8
2GLM-4.6V34
3Qwen 2.5 VL 7B Instruct33.1
4GPT-5.4 Mini32.1
5GPT-4o32
6InternVL3-8B31.2
7Qwen 3 VL 32B28.1
8Claude Sonnet 4.627.3
9InternVL3-38B26.2
10Claude Haiku 4.523.4

Interactive version: theaggregate.ai/benchmark?slug=roboprocessbench-current-primitive-recognition · How It Works · Data refreshed daily, snapshot 2026-09-29.