Chat2Workflow - Developer: leaderboard

Metric: Resolve rate (%) on Chat2Workflow developer tasks: share of test cases where the workflow the model generates for the current round of a multi-turn request, deployed on Dify 1.9.2, executes and its output satisfies every requirement of the instruction, judged by GPT-5.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)44.44
2GPT-5.238.89
3Claude Sonnet 4.531.94
4Kimi K2 (Thinking)26.39
5Qwen 3 235B A22B20.83
6GLM-4.719.45
7Qwen 3 32B15.28
8Qwen 3 14B11.11
9GPT-5.18.33
10Kimi K25.56
11GLM-4.65.55
12DeepSeek V3.24.17
13DeepSeek V3.12.78
14Qwen 3 8B1.39
15Qwen 3 Coder 480B A35B Instruct0

Interactive version: theaggregate.ai/benchmark?slug=chat2workflow-developer · How It Works · Data refreshed daily, snapshot 2026-10-07.