MUIAnno: leaderboard

Metric: F1-score (%) of UI element extraction on all 1,000 annotated iOS screens: a predicted element counts only if its box reaches IoU 0.5 with a ground-truth element under one-to-one matching and its type label matches; fixed prompt with one reference example, JSON-schema constrained output, deterministic API generation; times 100; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 5 models tracked.

Top models

#ModelScore
1GPT-5.470
2Claude Opus 4.667
3Gemini 3.1 Pro (Preview)55
4Llama 4 Scout44.5
5Gemma 4 31B (IT)41.6

Interactive version: theaggregate.ai/benchmark?slug=muianno · How It Works · Data refreshed daily, snapshot 2026-10-07.