Chat2Workflow - Enterprise: leaderboard

Metric: Resolve rate (%) on Chat2Workflow enterprise tasks: share of test cases where the workflow the model generates for the current round of a multi-turn request, deployed on Dify 1.9.2, executes and its output satisfies every requirement of the instruction, judged by GPT-5.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)64.82
2GPT-5.258.33
3GLM-4.744.45
4GPT-5.133.33
5Claude Sonnet 4.527.78
6Kimi K2 (Thinking)25.93
7DeepSeek V3.124.07
8GLM-4.624.07
9Kimi K28.33
10Qwen 3 32B5.55
11DeepSeek V3.21.85
12Qwen 3 235B A22B1.85
13Qwen 3 Coder 480B A35B Instruct1.85
14Qwen 3 8B0
15Qwen 3 14B0

Interactive version: theaggregate.ai/benchmark?slug=chat2workflow-enterprise · How It Works · Data refreshed daily, snapshot 2026-10-07.