CoCoBench - Sequential Ordering: leaderboard

Metric: Task success rate (%; sequential-ordering instances, respecting ordering constraints between agents; 897 oracle-validated executable household multi-agent task instances with two to four agents, centralized planner with image observations, task-specific step budgets). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.894.7
2GPT-5.6 Sol90.2
3Qwen 3.6 Plus61.8
4Claude Haiku 4.547.1
5Qwen 3.5 9B37.8
6GPT-5.4 Mini30.7
7Qwen 3 VL 8B Instruct4.9
8Kimi K2.63.6

Interactive version: theaggregate.ai/benchmark?slug=cocobench-sequential-ordering · How It Works · Data refreshed daily, snapshot 2026-09-26.