MobileDev-Bench (Oracle Retrieval): leaderboard
Metric: Overall resolution rate (%) with oracle retrieval (the ground-truth files to modify are given with the issue), MobileDev-Bench's 407 human-verified issue-resolution tasks from 19 production Android Native, React Native and Flutter apps; the model runs inside the Agentless localization-and-repair pipeline extended with tree-sitter parsing for Java, Kotlin, TypeScript and Dart, and a task counts as resolved when the generated patch passes the full test suite; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Agentless + Claude Sonnet 4.5 | 5.69 | |
| 2 | Agentless + GPT-5.2 | 4.53 | |
| 3 | Agentless + Qwen3-Coder [MobileDev-Bench checkpoint unspecified] | 2.57 | |
| 4 | Agentless + Gemini 2.5 Flash | 1.98 |
Interactive version: theaggregate.ai/benchmark?slug=mobiledev-bench-oracle-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-11.