ST-BiBench - Fine-Grained Action Control: leaderboard

Metric: Success rate (%) averaged over 5 atomic bimanual tasks in which the model directly outputs continuous 16-dimensional action streams (7-dimensional pose and 1-dimensional gripper per arm) from images plus ground-truth pose metadata, at least 100 episodes per task; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-566.8#91
2Gemini 2.5 Pro60.2#145
3Gemini 2.5 Flash53.6#237
4InternVL3-78B27.6#345
5Claude Sonnet 4.525.4#138
6Qwen 3 VL 235B A22B Instruct25.2#264
7Gemma 3 27B (IT)6.2#509
8Llama 4 Scout Instruct6#526

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=st-bibench-fine-grained-action-control · How It Works · Data refreshed daily, snapshot 2026-10-11.