AgentSocialBench - Task Completion Quality: leaderboard

Metric: Task Completion Quality (0-1; five-level grade of the coordination outcome against expert-annotated success criteria) averaged over scenarios, under the L0 (unconstrained) privacy instruction level, judged by Claude Opus 4.6; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.60.87
2Kimi K2.50.86
3Claude Sonnet 4.50.83
4DeepSeek V3.20.77
5MiniMax-M2.10.77
6Claude Haiku 4.50.73
7Qwen 3 235B A22B0.73
8GPT-5 Mini0.69

Interactive version: theaggregate.ai/benchmark?slug=agentsocialbench-task-completion-quality · How It Works · Data refreshed daily, snapshot 2026-10-07.