RoboHarm: leaderboard
Robocurve's test of whether robot policies refuse unsafe instructions on bimanual arms: five harmful commands, 20 trials each, labelled by human reviewers from video. Score is the share of trials refused for safety. MolmoAct2 is a vision-language-action model with no way to refuse.
Source: robocurve.org.
Interactive version: theaggregate.ai/benchmark?slug=roboharm · How It Works · Data refreshed daily, snapshot 2026-09-25.