AmbiBench - Requirement Coverage: leaderboard

Metric: Requirement coverage rate (%, 0-100): mean share of each task's atomic requirements (anchors, explicit constraints, implicit preferences) the agent satisfies; AmbiBench: 240 Android tasks in 25 real online apps (7 system, 18 third-party) at four instruction-clarity levels (detailed, standard, incomplete, ambiguous), run on real devices with a 25-step limit; a GPT-5 user simulator answers clarification requests only on incomplete and ambiguous tasks; the MUSE MLLM-as-a-judge pipeline audits each trajectory against the task's atomic requirements; average over the four clarity levels; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScoreOverall rank
1AutoGLM-9B (AmbiBench checkpoint unspecified)48.3
2MobileAgentV2 + GPT-4o33.8
3AppAgent + GPT-4o27.9
4UI-TARS-7B (AmbiBench checkpoint unspecified)24.6
5Qwen3-VL-8B (AmbiBench checkpoint unspecified)17.6

Interactive version: theaggregate.ai/benchmark?slug=ambibench-requirement-coverage · How It Works · Data refreshed daily, snapshot 2026-10-11.