AmbiBench: leaderboard

Metric: Task success rate (%, 0-100): share of tasks whose every atomic requirement is satisfied; AmbiBench: 240 Android tasks in 25 real online apps (7 system, 18 third-party) at four instruction-clarity levels (detailed, standard, incomplete, ambiguous), run on real devices with a 25-step limit; a GPT-5 user simulator answers clarification requests only on incomplete and ambiguous tasks; the MUSE MLLM-as-a-judge pipeline audits each trajectory against the task's atomic requirements; average over the four clarity levels; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScoreOverall rank
1AutoGLM-9B (AmbiBench checkpoint unspecified)35.4
2UI-TARS-7B (AmbiBench checkpoint unspecified)21.7
3MobileAgentV2 + GPT-4o20.4
4AppAgent + GPT-4o10.8
5Qwen3-VL-8B (AmbiBench checkpoint unspecified)10

Interactive version: theaggregate.ai/benchmark?slug=ambibench · How It Works · Data refreshed daily, snapshot 2026-10-11.