MuPPET - Utility: leaderboard

Metric: Utility (%): share of answers that use at least one useful preference or constraint extracted from the memories, judged by an LLM, in the main multi-party setting with no defence prompt: an LLM assistant acting for one user in a MuPPET multi-party workplace group conversation holds memory summaries with private information about that user and answers the group; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1GPT-5.582.03
2Qwen 3 14B81.14
3Qwen 3 8B78.29
4Gemini 2.5 Pro74.91
5Qwen 3 4B74.38
6Llama 3.1 8B Instruct69.93
7Llama 3 8B Instruct67.26

Interactive version: theaggregate.ai/benchmark?slug=muppet-utility · How It Works · Data refreshed daily, snapshot 2026-09-29.