MuPPET - One-on-One: leaderboard
Metric: Privacy leakage rate (%) in the one-on-one control setting: the model receives the full MuPPET item (multi-party workplace group conversation plus memory summaries with private information about the user it acts for) but is told that only the sender of the last message will see its reply, with a note asking it not to discuss the highly sensitive information with that recipient; a response leaks when a contextual-integrity LLM judge finds that it reveals the private detail directly or by strong implication; lower is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 9.88 |
| 2 | GPT-5.5 | 14.22 |
| 3 | Qwen 3 14B | 40.99 |
| 4 | Llama 3 8B Instruct | 45.68 |
| 5 | Qwen 3 4B | 53.09 |
| 6 | Llama 3.1 8B Instruct | 62.5 |
| 7 | Qwen 3 8B | 63.58 |
Interactive version: theaggregate.ai/benchmark?slug=muppet-one-on-one · How It Works · Data refreshed daily, snapshot 2026-09-29.