UXBench (Mobile) - Function Mismatch: leaderboard
Metric: Accuracy (0-100) on descriptions inconsistent with the provided functionality (trustworthiness): two- or three-option multiple-choice UX-defect diagnosis questions on real mobile app screenshots (2,000 questions over eight tasks, balanced positive and negative cases), temperature 0, 8,192-token output limit, a response that cannot be parsed counts as wrong; printed as a 0-1 accuracy and x100 here; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.7 Sonnet | 76 |
| 2 | Claude Sonnet 4 | 74 |
| 3 | Claude Sonnet 4.5 | 72 |
| 4 | Qwen 2.5 VL 72B Instruct | 71 |
| 5 | Qwen 3 VL 235B A22B Instruct | 70 |
| 6 | Qwen 3 VL 235B A22B (Thinking) | 70 |
| 7 | Qwen 3 VL 4B (Thinking) | 65 |
Interactive version: theaggregate.ai/benchmark?slug=uxbench-mobile-function-mismatch · How It Works · Data refreshed daily, snapshot 2026-09-29.