RoboGraphBench - Indoor: leaderboard

Metric: Task success rate (%; episodes that reach the goal and stop correctly, averaged over the baseline and six intervention conditions; semantic closed-loop evaluation on partially observed symbolic scene graphs, at most 100 steps; 189 navigation-enabled indoor episodes from 27 scenes). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1GPT-5.582.5
2Claude Opus 4.780.4
3GLM-5.257.1
4Gemini 3.1 Pro (Preview)42.3
5GPT-5.441.8
6Qwen 3.6 Plus28.6
7DeepSeek V4 Pro23.3
8Qwen 3.7 Plus21.7
9Qwen 3 VL 8B0.5

Interactive version: theaggregate.ai/benchmark?slug=robographbench-indoor · How It Works · Data refreshed daily, snapshot 2026-09-26.