ALERT — leaderboard
Large-scale red-teaming benchmark for assessing LLM safety across a structured risk taxonomy, with prompts designed to reveal weaknesses and unsafe behaviors.
Source: github.com.
Interactive version: theaggregate.ai/benchmark?slug=alert · How the rankings work · Data refreshed daily, snapshot 2026-07-22.