TRACE (LRM Safety) - Prompt Risk Category: leaderboard
Metric: Macro F1 (%; classification of unsafe prompts into nine risk categories by guardrails that support custom categories, TRACE English and Chinese). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gpt-oss-safeguard-20B | 62.25 |
Interactive version: theaggregate.ai/benchmark?slug=trace-lrm-safety-prompt-risk-category · How It Works · Data refreshed daily, snapshot 2026-09-26.