Chat2Workflow - AIGC: leaderboard

Metric: Resolve rate (%) on Chat2Workflow AI-generated content tasks: share of test cases where the workflow the model generates for the current round of a multi-turn request, deployed on Dify 1.9.2, executes and its output satisfies every requirement of the instruction, judged by GPT-5.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)54.32
2GLM-4.750.61
3GPT-5.145.06
4Kimi K2 (Thinking)43.83
5Claude Sonnet 4.541.97
6DeepSeek V3.140.12
7DeepSeek V3.239.51
8GLM-4.639.51
9Kimi K235.19
10GPT-5.232.1
11Qwen 3 32B29.63
12Qwen 3 235B A22B23.46
13Qwen 3 Coder 480B A35B Instruct23.46
14Qwen 3 14B22.22
15Qwen 3 8B6.79

Interactive version: theaggregate.ai/benchmark?slug=chat2workflow-aigc · How It Works · Data refreshed daily, snapshot 2026-10-07.