UI2App: leaderboard

Metric: Interaction Inference Score (0-100), the recall of the interactions the screenshots imply, each rated working, partial or failed by three annotators and weighted by state scope 1-3; 45 screenshot sets (327 screenshots) of real multi-route web applications; the model must emit a runnable React application from the screenshots alone; apps that fail to build after three self-debug rounds score 0. Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.639.3
2Kimi K2.5 (Thinking)20.7
3Qwen 3.5 397B A17B13.2
4Gemini 3.1 Pro (Preview)7.5
5GPT-5.46.7
6GLM-4.6V4.5

Interactive version: theaggregate.ai/benchmark?slug=ui2app · How It Works · Data refreshed daily, snapshot 2026-09-29.