UI2App - Executability (Self-Debug): leaderboard

Metric: EXEC@3 (%; share of applications that build and render after up to three rounds of feeding the build errors back to the model; 45 screenshot sets (327 screenshots) of real multi-route web applications; the model must emit a runnable React application from the screenshots alone). Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)100
2Claude Sonnet 4.6100
3Kimi K2.5 (Thinking)86.7
4GPT-5.482.2
5Qwen 3.5 397B A17B77.8
6GLM-4.6V35.6

Interactive version: theaggregate.ai/benchmark?slug=ui2app-executability-self-debug · How It Works · Data refreshed daily, snapshot 2026-09-29.