ROME Safety Judgment: leaderboard
Metric: Average F1 (%) of unsafe-trajectory detection over four balanced conditions: the 100 original R-Judge unsafe trajectories and their three ROME rewrites (implicit risks, contextual ambiguity, shortcut decision-making); the model judges zero-shot (temperature 0) whether an LLM-agent trajectory is safe or unsafe; unsafe is the positive class; each condition pairs 100 unsafe trajectories with the same 100 safe ones; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.7 Sonnet | 75.36 |
| 2 | Qwen 3 8B | 50.48 |
| 3 | GPT-4o (2024-11-20) | 43.27 |
Interactive version: theaggregate.ai/benchmark?slug=rome-safety-judgment · How It Works · Data refreshed daily, snapshot 2026-10-07.