PKU-SafeRLHF: leaderboard
Safety alignment benchmark/dataset separating helpfulness and harmlessness preferences, with harm categories and severity labels for multi-level risk control.
Source: github.com.
Interactive version: theaggregate.ai/benchmark?slug=pku-saferlhf · How It Works · Data refreshed daily, snapshot 2026-09-05.