Chat2Workflow - Education: leaderboard

Metric: Resolve rate (%) on Chat2Workflow education tasks: share of test cases where the workflow the model generates for the current round of a multi-turn request, deployed on Dify 1.9.2, executes and its output satisfies every requirement of the instruction, judged by GPT-5.1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.560.61
2Gemini 3 Pro (Preview)59.6
3GPT-5.254.55
4GLM-4.752.53
5DeepSeek V3.132.32
6Qwen 3 235B A22B29.29
7GLM-4.625.25
8Qwen 3 32B23.23
9DeepSeek V3.222.22
10GPT-5.119.19
11Qwen 3 Coder 480B A35B Instruct17.17
12Kimi K2 (Thinking)15.15
13Kimi K213.13
14Qwen 3 8B0
15Qwen 3 14B0

Interactive version: theaggregate.ai/benchmark?slug=chat2workflow-education · How It Works · Data refreshed daily, snapshot 2026-10-07.