Chat2Workflow: leaderboard

Metric: Resolve rate (%) on Chat2Workflow (79 instructions in 27 multi-turn tasks, three test cases each, pooled over six domains): share of test cases where the workflow the model generates for the current round of a multi-turn request, deployed on Dify 1.9.2, executes and its output satisfies every requirement of the instruction, judged by GPT-5.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)60.2
2Claude Sonnet 4.547.96
3GPT-5.247.4
4GLM-4.746.55
5Kimi K2 (Thinking)36
6GLM-4.635.02
7GPT-5.134.46
8DeepSeek V3.132.07
9DeepSeek V3.227.43
10Kimi K224.61
11Qwen 3 235B A22B24.47
12Qwen 3 32B19.13
13Qwen 3 Coder 480B A35B Instruct19.13
14Qwen 3 14B12.38
15Qwen 3 8B6.61

Interactive version: theaggregate.ai/benchmark?slug=chat2workflow · How It Works · Data refreshed daily, snapshot 2026-10-07.