MuPPET - Undefended Leakage: leaderboard

Metric: Privacy leakage rate (%) in the main multi-party setting with no defence prompt: an LLM assistant acting for one user in a MuPPET multi-party workplace group conversation holds memory summaries with private information about that user and answers the group; a response leaks when a contextual-integrity LLM judge finds that it reveals the private detail directly or by strong implication; lower is better. Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro39.14
2GPT-5.549.02
3Qwen 3 4B58.17
4Llama 3 8B Instruct58.29
5Qwen 3 14B64.53
6Llama 3.1 8B Instruct64.88
7Qwen 3 8B69.22

Interactive version: theaggregate.ai/benchmark?slug=muppet-undefended-leakage · How It Works · Data refreshed daily, snapshot 2026-09-29.