AgentSocialBench - Multi-Party Leakage: leaderboard

Metric: Privacy leakage rate (0-1, share of designated private items leaked partially or fully) of LLM agents in multi-party group chat and affinity-modulated scenarios (MPLR, over private item and recipient pairs), under the L0 (unconstrained) privacy instruction level, judged by Claude Opus 4.6; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.50.16
2Claude Sonnet 4.60.18
3GPT-5 Mini0.18
4MiniMax-M2.10.2
5DeepSeek V3.20.22
6Qwen 3 235B A22B0.22
7Claude Haiku 4.50.24
8Kimi K2.50.3

Interactive version: theaggregate.ai/benchmark?slug=agentsocialbench-multi-party-leakage · How It Works · Data refreshed daily, snapshot 2026-10-07.