CollabBench Cook - Game Score: leaderboard

Metric: Final game score of the two-agent kitchen team (completed dish orders), Cook-MultiPlayer (Overcooked-style cooperative cooking, five kitchen scenarios): the evaluated model drives Agent 1 inside the ProAgent framework while a DeepSeek-V3.1 simulator role-plays a partner with a Big Five personality profile; mean over all evaluation trajectories; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 4 models tracked.

Top models

#ModelScore
1DeepSeek V3.1136.53
2Qwen 2.5 72B Instruct135.47
3GPT-5.2135.2
4Qwen 2.5 7B Instruct86.93

Interactive version: theaggregate.ai/benchmark?slug=collabbench-cook-game-score · How It Works · Data refreshed daily, snapshot 2026-09-29.