AmbiBench - Requirement Coverage: leaderboard
Metric: Requirement coverage rate (%, 0-100): mean share of each task's atomic requirements (anchors, explicit constraints, implicit preferences) the agent satisfies; AmbiBench: 240 Android tasks in 25 real online apps (7 system, 18 third-party) at four instruction-clarity levels (detailed, standard, incomplete, ambiguous), run on real devices with a 25-step limit; a GPT-5 user simulator answers clarification requests only on incomplete and ambiguous tasks; the MUSE MLLM-as-a-judge pipeline audits each trajectory against the task's atomic requirements; average over the four clarity levels; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | AutoGLM-9B (AmbiBench checkpoint unspecified) | 48.3 | |
| 2 | MobileAgentV2 + GPT-4o | 33.8 | |
| 3 | AppAgent + GPT-4o | 27.9 | |
| 4 | UI-TARS-7B (AmbiBench checkpoint unspecified) | 24.6 | |
| 5 | Qwen3-VL-8B (AmbiBench checkpoint unspecified) | 17.6 |
Interactive version: theaggregate.ai/benchmark?slug=ambibench-requirement-coverage · How It Works · Data refreshed daily, snapshot 2026-10-11.