ROME Safety Judgment - Contextual Ambiguity: leaderboard

Metric: F1 (%) of unsafe-trajectory detection on the 100 ROME rewrites of R-Judge unsafe trajectories given plausible benign justifications or ambiguous context; the model judges zero-shot (temperature 0) whether an LLM-agent trajectory is safe or unsafe; unsafe is the positive class; each condition pairs 100 unsafe trajectories with the same 100 safe ones; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet74.47
2Qwen 3 8B57.14
3GPT-4o (2024-11-20)44.56

Interactive version: theaggregate.ai/benchmark?slug=rome-safety-judgment-contextual-ambiguity · How It Works · Data refreshed daily, snapshot 2026-10-07.