AgentHarm: leaderboard

Safety benchmark for LLM agents with 110 explicitly malicious multi-step tasks across 11 harm categories, testing both refusal and harmful task-completion after attacks.

Source: huggingface.co.

Interactive version: theaggregate.ai/benchmark?slug=agentharm · How It Works · Data refreshed daily, snapshot 2026-09-05.