CollabBench CWAH - Trustfulness: leaderboard

Metric: Trustfulness score (out of 5): trustfulness (reliable commitments and consistent, honest communication), rated by a penalty-based DeepSeek-V3.1 judge: each window of three consecutive agent actions starts at 5 and loses 1 to 3 points per violation, and the trajectory score subtracts violation and worst-window terms and a message-sparsity penalty, clipped at 0; CWAH-MultiPlayer (Communicative Watch-And-Help household tasks with two embodied agents, 10 held-out test scenarios): the evaluated model drives Agent 1 inside the CoELA agent framework while a DeepSeek-V3.1 simulator role-plays a partner with a Big Five personality profile; mean over all evaluation trajectories; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1GPT-5.23.74
2Qwen 2.5 72B Instruct3.61
3DeepSeek V3.13.35
4Qwen 2.5 7B Instruct2.58

Interactive version: theaggregate.ai/benchmark?slug=collabbench-cwah-trustfulness · How It Works · Data refreshed daily, snapshot 2026-09-29.