UXBench (Mobile) - Content Mismatch: leaderboard

Metric: Accuracy (0-100) on service names inconsistent with the page text (trustworthiness): two- or three-option multiple-choice UX-defect diagnosis questions on real mobile app screenshots (2,000 questions over eight tasks, balanced positive and negative cases), temperature 0, 8,192-token output limit, a response that cannot be parsed counts as wrong; printed as a 0-1 accuracy and x100 here; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 14 models tracked.

Top models

#ModelScore
1Qwen 2.5 VL 72B Instruct68
2Claude Sonnet 4.566
3Claude Sonnet 466
4Claude 3.7 Sonnet65
5Qwen 3 VL 235B A22B Instruct64
6Qwen 3 VL 235B A22B (Thinking)64
7Qwen 3 VL 4B (Thinking)63

Interactive version: theaggregate.ai/benchmark?slug=uxbench-mobile-content-mismatch · How It Works · Data refreshed daily, snapshot 2026-09-29.