UI2App: leaderboard
Metric: Interaction Inference Score (0-100), the recall of the interactions the screenshots imply, each rated working, partial or failed by three annotators and weighted by state scope 1-3; 45 screenshot sets (327 screenshots) of real multi-route web applications; the model must emit a runnable React application from the screenshots alone; apps that fail to build after three self-debug rounds score 0. Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 39.3 |
| 2 | Kimi K2.5 (Thinking) | 20.7 |
| 3 | Qwen 3.5 397B A17B | 13.2 |
| 4 | Gemini 3.1 Pro (Preview) | 7.5 |
| 5 | GPT-5.4 | 6.7 |
| 6 | GLM-4.6V | 4.5 |
Interactive version: theaggregate.ai/benchmark?slug=ui2app · How It Works · Data refreshed daily, snapshot 2026-09-29.