GMA - Atomic: leaderboard

Metric: Success rate (%; 300 tasks in seven custom Android applications, vision-only agent seeing raw screenshots (two most recent kept), MobileWorld action space plus call_user and answer, Qwen3.7-Plus user simulator, 150-step budget; 73 atomic single-action tasks). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1Qwen 3.8 Max93.15
2Seed 2.1 Pro89.04
3GPT-5.587.67
4Claude Opus 4.787.67
5GPT-5.6 Sol86.3
6Claude Sonnet 4.683.56
7Gemini 3.5 Flash80.82
8Qwen 3.7 Plus78.08

Interactive version: theaggregate.ai/benchmark?slug=gma-atomic · How It Works · Data refreshed daily, snapshot 2026-09-29.