CooperBench: leaderboard

Metric: Cooperative task success (%; two agents that can message each other each implement one feature of a task pair in isolated workspaces, and the merged patch must pass both features' tests after the naive, union and learned-resolver merge stages; 652 tasks). Source: arxiv.org. Saturation forecast: Around October 2027. 5 models tracked.

Top models

#ModelScore
1GPT-527.9
2Claude Sonnet 4.525.92
3MiniMax-M213.96
4Qwen 3 Coder 30B A3B Instruct13.34
5Qwen 3 30B A3B 2507 Instruct4.6

Interactive version: theaggregate.ai/benchmark?slug=cooperbench · How It Works · Data refreshed daily, snapshot 2026-09-26.