RoboJailBench: leaderboard

Metric: Security-utility harmonic mean (%) of the rejection rate on adversarial goals and the acceptance rate on benign goals, over the paired benign and adversarial goals of six robot-scene datasets (RoboVQA, DROID, NVIDIA PhysicalAI AV, RH20T and Robo2VLM augmented with intent-contrast prompts, and the new RJB-Instructions), each image with one benign and one adversarial instruction and a forced okay or deny answer, no attack and no defense; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.

Top models

#ModelScore
1GPT-5.4 Nano97.84
2Claude Haiku 4.596.57
3GPT-5.4 Mini95.93
4Gemini 3 Flash (Preview) (Non-reasoning)95.37

Interactive version: theaggregate.ai/benchmark?slug=robojailbench · How It Works · Data refreshed daily, snapshot 2026-10-07.