AgentHarm — leaderboard
Safety benchmark for LLM agents with 110 explicitly malicious multi-step tasks across 11 harm categories, testing both refusal and harmful task-completion after attacks.
Source: huggingface.co.
Interactive version: theaggregate.ai/benchmark?slug=agentharm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.