UXBench (Mobile) - Popup Stacking: leaderboard

Metric: Accuracy (0-100) on several modal popups shown at once (efficiency): two- or three-option multiple-choice UX-defect diagnosis questions on real mobile app screenshots (2,000 questions over eight tasks, balanced positive and negative cases), temperature 0, 8,192-token output limit, a response that cannot be parsed counts as wrong; printed as a 0-1 accuracy and x100 here; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.566
2Qwen 2.5 VL 72B Instruct56
3Qwen 3 VL 235B A22B (Thinking)56
4Claude 3.7 Sonnet53
5Claude Sonnet 452
6Qwen 3 VL 235B A22B Instruct52
7Qwen 3 VL 4B (Thinking)50
8GLM-4.1V-9B (Thinking)32

Interactive version: theaggregate.ai/benchmark?slug=uxbench-mobile-popup-stacking · How It Works · Data refreshed daily, snapshot 2026-09-29.