EvoGenUI-Bench - Episode Pass: leaderboard
Metric: Episode pass rate (%; mean of the three suites' shares of five-turn episodes in which all five turns pass; 150 five-turn generative-UI tasks (750 turns) in three suites of 50: information presentation, executable interaction and tool-grounded external state; each turn's artifact is executed in a browser and a MiMo-V2.5 evaluator checks hidden requirements from screenshots, source and DOM evidence, MiMo-V2.5 actor traces and runtime logs; turns left unexecuted after an earlier build failure or invalid output count as failures; temperature 0). Source: arxiv.org. Saturation forecast: Around December 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.7 | 37.3 |
| 2 | GPT-5.5 | 21.3 |
| 3 | Claude Haiku 4.5 (20251001) | 13.3 |
| 4 | Qwen 3.6 Plus | 7.3 |
| 5 | Gemini 3.1 Pro (Preview) | 6.7 |
| 6 | Gemini 3 Flash (Preview) | 4.7 |
| 7 | Qwen 3 Coder 480B A35B Instruct | 2 |
| 8 | GLM 4.5 Air | 1.3 |
Interactive version: theaggregate.ai/benchmark?slug=evogenui-bench-episode-pass · How It Works · Data refreshed daily, snapshot 2026-09-29.