TRACE (LRM Safety) - Prompt: leaderboard

Metric: F1 on the unsafe class (%; guardrail safe or unsafe judgments of the prompt in 5,000 prompt, reasoning-trace and final-response triples from four reasoning models, English and Chinese, nine risk categories and ten attack strategies; labels by a majority of three LLM annotators, temperature 0). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1gpt-oss-safeguard-20B86.96

Interactive version: theaggregate.ai/benchmark?slug=trace-lrm-safety-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.