ST-BiBench - Foundational Spatial Grounding: leaderboard

Metric: Success score (%) for assigning the reachable arm and target in bimanual scenes, the mean of the sparse, dense and cluttered (with distractors) scenario settings, at least 100 domain-randomized episodes each; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 29 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.0 Flash95.38#331
2Gemini 2.5 Flash95.13#237
3Gemini 2.5 Pro95.01#145
4Claude Sonnet 4.594.38#138
5GPT-594.28#91
6GLM-4.5V94.08#339
7Qwen 3 VL 32B Instruct94#276
8Claude Sonnet 4 (20250514)93.82#211
9Claude 3.7 Sonnet (20250219)93.51#196
10InternVL3-78B93.34#345
11GPT-4.192.55#240
12Qwen 3 VL 235B A22B Instruct90.22#264
13GPT-4o89.08#333
14Qwen 3 VL 30B A3B Instruct88.54#365
15InternVL3-38B87.94#395

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=st-bibench-foundational-spatial-grounding · How It Works · Data refreshed daily, snapshot 2026-10-11.