WildGuardTest — leaderboard
Human-annotated safety moderation benchmark for harmful prompt detection, response harmfulness evaluation, and refusal identification, including jailbreak-style cases.
Source: huggingface.co.
Interactive version: theaggregate.ai/benchmark?slug=wildguardtest · How the rankings work · Data refreshed daily, snapshot 2026-07-22.