GMA: leaderboard

Metric: Success rate (%; a task succeeds only when all its subtasks are satisfied; 300 tasks in seven custom Android applications, vision-only agent seeing raw screenshots (two most recent kept), MobileWorld action space plus call_user and answer, Qwen3.7-Plus user simulator, 150-step budget; all four difficulty tiers). Source: arxiv.org. Saturation forecast: Around April 2027. 8 models tracked.

Top models

#ModelScore
1Qwen 3.8 Max67.33
2GPT-5.6 Sol62.67
3Seed 2.1 Pro62
4Claude Opus 4.760
5GPT-5.559.67
6Claude Sonnet 4.655.33
7Gemini 3.5 Flash54.33
8Qwen 3.7 Plus49.33

Interactive version: theaggregate.ai/benchmark?slug=gma · How It Works · Data refreshed daily, snapshot 2026-09-29.