GTA-Workflow - Root Score: leaderboard

Metric: Mean root checkpoint score (0-10) on GTA-Workflow, the open-ended long-horizon workflow tier of GTA-2 (132 real-user tasks with deployed tools and multimodal inputs), run end to end in the default Lagent ReAct framework; deliverables judged by GPT-5.2 against recursive checkpoint rubrics; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 13 models tracked.

Top models

#ModelScore
1GPT-53.66
2Llama 4 Scout3.65
3Gemini 2.5 Pro3.64
4Qwen 3 235B A22B3.59
5DeepSeek V3.23.56
6Grok 43.56
7Claude Sonnet 4.53.5
8Kimi K23.5
9Qwen 3 8B1.81
10Llama 3.1 70B Instruct1.55
11Qwen 3 30B A3B1.21
12Llama 3.1 8B Instruct1.18
13Llama 3.2 3B Instruct1.02

Interactive version: theaggregate.ai/benchmark?slug=gta-workflow-root-score · How It Works · Data refreshed daily, snapshot 2026-10-07.