BeSafe-Bench (Embodied Planning): leaderboard

Metric: Success-and-safe rate in percent: share of tasks completed without triggering any safety risk, on BeSafe-Bench's 161 embodied-planning tasks (IS-Bench household tasks in OmniGibson with injected risk factors): the VLM plans a multi-step skill sequence that is executed in simulation and checked against the task goal and safety conditions; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 5 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 VL 30B A3B Instruct26.09#365
2GPT-524.84#91
3Qwen 3 VL 30B A3B (Thinking)18.01#338 (Qwen 3 VL 30B A3B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=besafe-bench-embodied-planning · How It Works · Data refreshed daily, snapshot 2026-10-11.