MobileForge - Visual Fidelity: leaderboard

Metric: Point-wise visual fidelity (1-5): mean four-dimension rubric score a Gemini 2.5 Pro vision-language judge gives each generated screen against its reference design; 29 real mobile apps given as 309 screenshots with page-relationship annotations, each turned into a React, TypeScript and Tailwind project by one tool-using agent in a fixed harness (up to 50 iterations, build errors fed back); one run per app; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.62.99
2GPT-52.64
3Gemini 2.5 Pro2.56
4Claude Haiku 4.52.37
5Gemini 2.5 Flash2
6GPT-5 Mini1.6

Interactive version: theaggregate.ai/benchmark?slug=mobileforge-visual-fidelity · How It Works · Data refreshed daily, snapshot 2026-09-29.