ST-BiBench - Sequential Collaborative Manipulation: leaderboard

Metric: Success rate (%) over the 8 sequential collaborative manipulation tasks (hand-overs and multi-stage placements where one arm's step depends on the other's), each run for at least 100 domain-randomized episodes in which the model plans a sequence of seven parameterized bimanual action primitives executed through an API, averaged over the tasks it was scored on; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 28 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro69.38#145
2GPT-559.75#91
3Gemini 2.5 Flash59#237
4Qwen 3 VL 32B Instruct50.88#276
5Qwen 3 VL 235B A22B Instruct50.88#264
6Qwen 2.5 VL 32B Instruct50.13#443
7InternVL3-38B49.38#395
8Claude Sonnet 4 (20250514)46.63#211
9GPT-4o45.5#333
10Claude 3.7 Sonnet (20250219)45#196
11GPT-4.142.88#240
12Gemini 2.0 Flash41.25#331
13Qwen 2.5 VL 72B Instruct37.25#364
14InternVL3-78B33.63#345
15Llama 4 Scout Instruct29.75#526

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=st-bibench-sequential-collaborative-manipulation · How It Works · Data refreshed daily, snapshot 2026-10-11.