MobileDev-Bench: leaderboard

Metric: Overall resolution rate (%) with automated retrieval (the pipeline must localize the files to edit from the issue alone; instances without a submitted patch are left out of the rate), MobileDev-Bench's 407 human-verified issue-resolution tasks from 19 production Android Native, React Native and Flutter apps; the model runs inside the Agentless localization-and-repair pipeline extended with tree-sitter parsing for Java, Kotlin, TypeScript and Dart, and a task counts as resolved when the generated patch passes the full test suite; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScoreOverall rank
1Agentless + Claude Sonnet 4.54.23
2Agentless + Qwen3-Coder [MobileDev-Bench checkpoint unspecified]3.33
3Agentless + Gemini 2.5 Flash3.24
4Agentless + GPT-5.23.23

Interactive version: theaggregate.ai/benchmark?slug=mobiledev-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.