AgentHarm: leaderboard
Safety benchmark for LLM agents with 110 explicitly malicious multi-step tasks across 11 harm categories, testing both refusal and harmful task-completion after attacks.
Source: huggingface.co.
Interactive version: theaggregate.ai/benchmark?slug=agentharm · How It Works · Data refreshed daily, snapshot 2026-09-05.