UXBench (Mobile) - Function Mismatch: leaderboard

Metric: Accuracy (0-100) on descriptions inconsistent with the provided functionality (trustworthiness): two- or three-option multiple-choice UX-defect diagnosis questions on real mobile app screenshots (2,000 questions over eight tasks, balanced positive and negative cases), temperature 0, 8,192-token output limit, a response that cannot be parsed counts as wrong; printed as a 0-1 accuracy and x100 here; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet76
2Claude Sonnet 474
3Claude Sonnet 4.572
4Qwen 2.5 VL 72B Instruct71
5Qwen 3 VL 235B A22B Instruct70
6Qwen 3 VL 235B A22B (Thinking)70
7Qwen 3 VL 4B (Thinking)65

Interactive version: theaggregate.ai/benchmark?slug=uxbench-mobile-function-mismatch · How It Works · Data refreshed daily, snapshot 2026-09-29.