RoboJailBench - Utility Rate: leaderboard
Metric: Utility rate (%), the share of benign goals accepted, over the paired benign and adversarial goals of six robot-scene datasets (RoboVQA, DROID, NVIDIA PhysicalAI AV, RH20T and Robo2VLM augmented with intent-contrast prompts, and the new RJB-Instructions), each image with one benign and one adversarial instruction and a forced okay or deny answer, no attack and no defense; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 Nano | 99.49 |
| 2 | Gemini 3 Flash (Preview) (Non-reasoning) | 97.09 |
| 3 | GPT-5.4 Mini | 95.87 |
| 4 | Claude Haiku 4.5 | 94.91 |
Interactive version: theaggregate.ai/benchmark?slug=robojailbench-utility-rate · How It Works · Data refreshed daily, snapshot 2026-10-07.