EvoGenUI-Bench - Tool-Grounded: leaderboard

Metric: Turn pass rate (%; share of the 250 tool-grounded external-state turns whose artifact passes every hidden requirement; 150 five-turn generative-UI tasks (750 turns) in three suites of 50: information presentation, executable interaction and tool-grounded external state; each turn's artifact is executed in a browser and a MiMo-V2.5 evaluator checks hidden requirements from screenshots, source and DOM evidence, MiMo-V2.5 actor traces and runtime logs; turns left unexecuted after an earlier build failure or invalid output count as failures; temperature 0). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1GPT-5.562.8
2Claude Opus 4.761.2
3Qwen 3.6 Plus25.2
4Claude Haiku 4.5 (20251001)21.2
5Gemini 3 Flash (Preview)14.8
6Gemini 3.1 Pro (Preview)7.2
7Qwen 3 Coder 480B A35B Instruct4.4
8GLM 4.5 Air3.2

Interactive version: theaggregate.ai/benchmark?slug=evogenui-bench-tool-grounded · How It Works · Data refreshed daily, snapshot 2026-09-29.