AgentSocialBench - Hub-and-Spoke Leakage: leaderboard

Metric: Privacy leakage rate (0-1, share of designated private items leaked partially or fully) of LLM agents in hub-and-spoke scenarios (HALR, cross-contamination between participant pairs through a coordinator), under the L0 (unconstrained) privacy instruction level, judged by Claude Opus 4.6; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Qwen 3 235B A22B0.06
2Claude Sonnet 4.60.1
3Claude Sonnet 4.50.1
4GPT-5 Mini0.11
5Kimi K2.50.12
6DeepSeek V3.20.14
7Claude Haiku 4.50.15
8MiniMax-M2.10.2

Interactive version: theaggregate.ai/benchmark?slug=agentsocialbench-hub-and-spoke-leakage · How It Works · Data refreshed daily, snapshot 2026-10-07.