AgentDojo: leaderboard

Dynamic benchmark for prompt-injection attacks and defenses in tool-using LLM agents, with realistic tasks and security test cases across banking, Slack, travel, and workspace suites.

Source: github.com.

Interactive version: theaggregate.ai/benchmark?slug=agentdojo · How It Works · Data refreshed daily, snapshot 2026-09-05.