RoboJailBench: leaderboard
Metric: Security-utility harmonic mean (%) of the rejection rate on adversarial goals and the acceptance rate on benign goals, over the paired benign and adversarial goals of six robot-scene datasets (RoboVQA, DROID, NVIDIA PhysicalAI AV, RH20T and Robo2VLM augmented with intent-contrast prompts, and the new RJB-Instructions), each image with one benign and one adversarial instruction and a forced okay or deny answer, no attack and no defense; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 Nano | 97.84 |
| 2 | Claude Haiku 4.5 | 96.57 |
| 3 | GPT-5.4 Mini | 95.93 |
| 4 | Gemini 3 Flash (Preview) (Non-reasoning) | 95.37 |
Interactive version: theaggregate.ai/benchmark?slug=robojailbench · How It Works · Data refreshed daily, snapshot 2026-10-07.