CollabBench Cook - Helpfulness: leaderboard
Metric: Helpfulness score (out of 5): helpfulness (task focus, proactive support and useful coordination), rated by a penalty-based DeepSeek-V3.1 judge: each window of three consecutive agent actions starts at 5 and loses 1 to 3 points per violation, and the trajectory score subtracts violation and worst-window terms and a message-sparsity penalty, clipped at 0; Cook-MultiPlayer (Overcooked-style cooperative cooking, five kitchen scenarios): the evaluated model drives Agent 1 inside the ProAgent framework while a DeepSeek-V3.1 simulator role-plays a partner with a Big Five personality profile; mean over all evaluation trajectories; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3.1 | 1.79 |
| 2 | GPT-5.2 | 1.63 |
| 3 | Qwen 2.5 72B Instruct | 1.37 |
| 4 | Qwen 2.5 7B Instruct | 0.45 |
Interactive version: theaggregate.ai/benchmark?slug=collabbench-cook-helpfulness · How It Works · Data refreshed daily, snapshot 2026-09-29.