TRACE (LRM Safety) - Reasoning Trace: leaderboard
Metric: F1 on the unsafe class (%; guardrail safe or unsafe judgments of the reasoning trace in 5,000 prompt, reasoning-trace and final-response triples from four reasoning models, English and Chinese, nine risk categories and ten attack strategies; labels by a majority of three LLM annotators, temperature 0). Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gpt-oss-safeguard-20B | 79.6 |
Interactive version: theaggregate.ai/benchmark?slug=trace-lrm-safety-reasoning-trace · How It Works · Data refreshed daily, snapshot 2026-09-26.