GMA - Compositional: leaderboard

Metric: Success rate (%; 300 tasks in seven custom Android applications, vision-only agent seeing raw screenshots (two most recent kept), MobileWorld action space plus call_user and answer, Qwen3.7-Plus user simulator, 150-step budget; 129 compositional multi-step tasks in one application). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Qwen 3.8 Max70.54
2GPT-5.6 Sol68.99
3Seed 2.1 Pro65.89
4GPT-5.565.12
5Claude Opus 4.765.12
6Gemini 3.5 Flash62.02
7Claude Sonnet 4.660.47
8Qwen 3.7 Plus51.94

Interactive version: theaggregate.ai/benchmark?slug=gma-compositional · How It Works · Data refreshed daily, snapshot 2026-09-29.