HarmBench: leaderboard

Standardized framework for automated red teaming and refusal testing. Evaluates attacks and defenses across harmful behavior categories and model families.

Source: github.com.

Interactive version: theaggregate.ai/benchmark?slug=harmbench · How It Works · Data refreshed daily, snapshot 2026-09-05.