GTA-Workflow - Perception: leaderboard

Metric: Root success rate (%) on the tasks that require perception tools on GTA-Workflow, the open-ended long-horizon workflow tier of GTA-2 (132 real-user tasks with deployed tools and multimodal inputs), run end to end in the default Lagent ReAct framework; deliverables judged by GPT-5.2 against recursive checkpoint rubrics; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 13 models tracked.

Top models

#ModelScore
1Qwen 3 235B A22B15.79
2Llama 4 Scout15.79
3Gemini 2.5 Pro13.16
4GPT-513.16
5Claude Sonnet 4.510.53
6DeepSeek V3.210.53
7Kimi K210.53
8Grok 47.89
9Llama 3.1 70B Instruct2.63
10Qwen 3 30B A3B2.63
11Llama 3.1 8B Instruct0
12Qwen 3 8B0
13Llama 3.2 3B Instruct0

Interactive version: theaggregate.ai/benchmark?slug=gta-workflow-perception · How It Works · Data refreshed daily, snapshot 2026-10-07.