AmbiBench: leaderboard
Metric: Task success rate (%, 0-100): share of tasks whose every atomic requirement is satisfied; AmbiBench: 240 Android tasks in 25 real online apps (7 system, 18 third-party) at four instruction-clarity levels (detailed, standard, incomplete, ambiguous), run on real devices with a 25-step limit; a GPT-5 user simulator answers clarification requests only on incomplete and ambiguous tasks; the MUSE MLLM-as-a-judge pipeline audits each trajectory against the task's atomic requirements; average over the four clarity levels; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | AutoGLM-9B (AmbiBench checkpoint unspecified) | 35.4 | |
| 2 | UI-TARS-7B (AmbiBench checkpoint unspecified) | 21.7 | |
| 3 | MobileAgentV2 + GPT-4o | 20.4 | |
| 4 | AppAgent + GPT-4o | 10.8 | |
| 5 | Qwen3-VL-8B (AmbiBench checkpoint unspecified) | 10 |
Interactive version: theaggregate.ai/benchmark?slug=ambibench · How It Works · Data refreshed daily, snapshot 2026-10-11.