CollabBench CWAH - Completion Steps: leaderboard

Metric: Task completion steps: environment steps the two-agent team needs to finish the household goal, CWAH-MultiPlayer (Communicative Watch-And-Help household tasks with two embodied agents, 10 held-out test scenarios): the evaluated model drives Agent 1 inside the CoELA agent framework while a DeepSeek-V3.1 simulator role-plays a partner with a Big Five personality profile; mean over all evaluation trajectories; lower is better. Source: arxiv.org. Saturation forecast: Around April 2027. 4 models tracked.

Top models

#ModelScore
1GPT-5.267.49
2Qwen 2.5 72B Instruct68.68
3DeepSeek V3.169.26
4Qwen 2.5 7B Instruct84.51

Interactive version: theaggregate.ai/benchmark?slug=collabbench-cwah-completion-steps · How It Works · Data refreshed daily, snapshot 2026-09-29.