MobilePA-Bench - Sub-agent Collaboration: leaderboard

Metric: Routing-and-handoff joint success (%). Source: arxiv.org. 13 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)77.53
2Claude Fable 570.79
3Gemini 3.6 Flash66.29
4Kimi K362.92
5Claude Opus 562.92
6GLM-5.261.8
7Seed 2.1 Pro59.55
8Qwen 3.8 Max53.93
9GPT-5.551.69
10Claude Opus 4.850.56
11Qwen 3.7 Max50.56
12GPT-5.6 Sol49.44
13Kimi K2.643.82

Interactive version: theaggregate.ai/benchmark?slug=mobilepa-bench-sub-agent-collaboration · How It Works · Data refreshed daily, snapshot 2026-09-19.