EntCollabBench - Workflow: leaderboard

Metric: Task success (%) on the 200 operational workflow tasks (160 single-step and 40 multi-step), judged from execution traces and service-state diffs, role-specialized LLM agents (11 enterprise roles behind permission-controlled ITSM, HR, CSM, Gitea and document services) delegate over HTTP, temperature 0, three-judge majority vote; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro59
2DeepSeek V4 Flash51
3Claude Sonnet 4.644
4Gemini 3.1 Pro (Preview)42.5
5GPT-5.433.5
6Qwen 3.5 122B A10B27.5
7Qwen 3.5 35B A3B25
8MiMo-V2-Flash21.5
9MiniMax-M2.718.5
10GPT-5 Mini17
11Gemini 3.1 Flash Lite (Preview)2
12Qwen 3.5 9B0

Interactive version: theaggregate.ai/benchmark?slug=entcollabbench-workflow · How It Works · Data refreshed daily, snapshot 2026-10-07.