ST-BiBench - Independent Parallel Manipulation: leaderboard

Metric: Success rate (%) over the 6 independent parallel manipulation tasks (place, rank and stack with both arms at once), each run for at least 100 domain-randomized episodes in which the model plans a sequence of seven parameterized bimanual action primitives executed through an API, averaged over the tasks it was scored on; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 28 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4.178.5#240
2GPT-576.67#91
3Gemini 2.5 Pro71.33#145
4Claude 3.7 Sonnet (20250219)69#196
5Gemini 2.5 Flash67.17#237
6Claude Sonnet 4 (20250514)67#211
7Gemini 2.0 Flash62.83#331
8Qwen 3 VL 235B A22B Instruct58.67#264
9InternVL3-38B57.5#395
10InternVL3-78B56.33#345
11Qwen 3 VL 32B Instruct54.67#276
12Qwen 2.5 VL 32B Instruct52.67#443
13GPT-4o52.33#333
14InternVL2.5-78B47.83#421
15Qwen 2.5 VL 72B Instruct28.6#364

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=st-bibench-independent-parallel-manipulation · How It Works · Data refreshed daily, snapshot 2026-10-11.