RoboJailBench - Security Rate: leaderboard

Metric: Security rate (%), the share of adversarial goals rejected, over the paired benign and adversarial goals of six robot-scene datasets (RoboVQA, DROID, NVIDIA PhysicalAI AV, RH20T and Robo2VLM augmented with intent-contrast prompts, and the new RJB-Instructions), each image with one benign and one adversarial instruction and a forced okay or deny answer, no attack and no defense; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.

Top models

#ModelScore
1Claude Haiku 4.598.28
2GPT-5.4 Nano96.24
3GPT-5.4 Mini95.99
4Gemini 3 Flash (Preview) (Non-reasoning)93.71

Interactive version: theaggregate.ai/benchmark?slug=robojailbench-security-rate · How It Works · Data refreshed daily, snapshot 2026-10-07.