PrepBench - GUI Workflow: leaderboard

Metric: End-to-end GUI workflow accuracy (%): from the original ambiguous request, the agent must produce an executable operator workflow whose output table matches the reference, PrepBench natural-language data-preparation tasks, LLM agent at temperature 0.7 with Clarify (budgeted, answered by a DeepSeek-V3.2 user simulator), Profile, Code and Translate actions; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 10 models tracked.

Top models

#ModelScore
1GPT-5.1 Codex34.6
2Kimi K2 (Thinking)30.1
3Claude Sonnet 4.524.5
4Gemini 3 Flash22.2
5GLM-4.719.9
6Qwen 3 235B A22B19.3
7DeepSeek V3.215.7
8Grok Code Fast 113.1
9Devstral 28.5
10GPT-4o5.2

Interactive version: theaggregate.ai/benchmark?slug=prepbench-gui-workflow · How It Works · Data refreshed daily, snapshot 2026-10-07.