MobileDev-Bench: leaderboard
Metric: Overall resolution rate (%) with automated retrieval (the pipeline must localize the files to edit from the issue alone; instances without a submitted patch are left out of the rate), MobileDev-Bench's 407 human-verified issue-resolution tasks from 19 production Android Native, React Native and Flutter apps; the model runs inside the Agentless localization-and-repair pipeline extended with tree-sitter parsing for Java, Kotlin, TypeScript and Dart, and a task counts as resolved when the generated patch passes the full test suite; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Agentless + Claude Sonnet 4.5 | 4.23 | |
| 2 | Agentless + Qwen3-Coder [MobileDev-Bench checkpoint unspecified] | 3.33 | |
| 3 | Agentless + Gemini 2.5 Flash | 3.24 | |
| 4 | Agentless + GPT-5.2 | 3.23 |
Interactive version: theaggregate.ai/benchmark?slug=mobiledev-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.