AgentLeak: leaderboard

Full-stack privacy leakage benchmark for multi-agent LLM systems, auditing internal channels such as inter-agent messages, shared memory, and tool arguments in addition to final outputs.

Metric: Total Leak (self-reported). Source: benchmarklist.com. 5 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet55.2
2GPT-4o Mini76.3
3GPT-4o77.6
4Llama 3.3 70B89.9
5Mistral Large99.3

Interactive version: theaggregate.ai/benchmark?slug=agentleak · How It Works · Data refreshed daily, snapshot 2026-09-05.