ST-BiBench - Strategic Coordination Planning: leaderboard

Metric: Success rate (%) over the 14 strategic coordination tasks (6 independent parallel and 8 sequential collaborative manipulation tasks), each run for at least 100 domain-randomized episodes in which the model plans a sequence of seven parameterized bimanual action primitives executed through an API, averaged over the tasks it was scored on; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 28 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro70.21#145
2GPT-567#91
3Gemini 2.5 Flash62.5#237
4GPT-4.158.14#240
5Claude Sonnet 4 (20250514)55.36#211
6Claude 3.7 Sonnet (20250219)55.29#196
7Qwen 3 VL 235B A22B Instruct54.21#264
8InternVL3-38B52.86#395
9Qwen 3 VL 32B Instruct52.5#276
10Qwen 2.5 VL 32B Instruct51.21#443
11Gemini 2.0 Flash50.5#331
12GPT-4o48.43#333
13InternVL3-78B43.36#345
14InternVL2.5-78B37.36#421
15Qwen 2.5 VL 72B Instruct33.92#364

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=st-bibench-strategic-coordination-planning · How It Works · Data refreshed daily, snapshot 2026-10-11.