MobileDev-Bench (Oracle Retrieval): leaderboard

Metric: Overall resolution rate (%) with oracle retrieval (the ground-truth files to modify are given with the issue), MobileDev-Bench's 407 human-verified issue-resolution tasks from 19 production Android Native, React Native and Flutter apps; the model runs inside the Agentless localization-and-repair pipeline extended with tree-sitter parsing for Java, Kotlin, TypeScript and Dart, and a task counts as resolved when the generated patch passes the full test suite; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScoreOverall rank
1Agentless + Claude Sonnet 4.55.69
2Agentless + GPT-5.24.53
3Agentless + Qwen3-Coder [MobileDev-Bench checkpoint unspecified]2.57
4Agentless + Gemini 2.5 Flash1.98

Interactive version: theaggregate.ai/benchmark?slug=mobiledev-bench-oracle-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-11.