AgentLeak — leaderboard

Full-stack privacy leakage benchmark for multi-agent LLM systems, auditing internal channels such as inter-agent messages, shared memory, and tool arguments in addition to final outputs.

Metric: Total Leak (self-reported). Source: benchmarklist.com. 5 models tracked.

Top models

#ModelScore
1Mistral Large99.3
2Llama 3.3 70B Instruct89.9
3GPT-4o77.6
4GPT-4o Mini76.3
5Claude 3.5 Sonnet55.2

Interactive version: theaggregate.ai/benchmark?slug=agentleak · How the rankings work · Data refreshed daily, snapshot 2026-07-22.