HarmBench — leaderboard
Standardized framework for automated red teaming and robust refusal. Evaluates attacks and defenses across harmful behavior categories and model families.
Source: github.com.
Interactive version: theaggregate.ai/benchmark?slug=harmbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.