UXBench (Mobile) - Content Mismatch: leaderboard
Metric: Accuracy (0-100) on service names inconsistent with the page text (trustworthiness): two- or three-option multiple-choice UX-defect diagnosis questions on real mobile app screenshots (2,000 questions over eight tasks, balanced positive and negative cases), temperature 0, 8,192-token output limit, a response that cannot be parsed counts as wrong; printed as a 0-1 accuracy and x100 here; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 VL 72B Instruct | 68 |
| 2 | Claude Sonnet 4.5 | 66 |
| 3 | Claude Sonnet 4 | 66 |
| 4 | Claude 3.7 Sonnet | 65 |
| 5 | Qwen 3 VL 235B A22B Instruct | 64 |
| 6 | Qwen 3 VL 235B A22B (Thinking) | 64 |
| 7 | Qwen 3 VL 4B (Thinking) | 63 |
Interactive version: theaggregate.ai/benchmark?slug=uxbench-mobile-content-mismatch · How It Works · Data refreshed daily, snapshot 2026-09-29.