WebIGBench (Action Prompt): leaderboard
Metric: Success rate (%; the UI agent replays the reference interaction on the generated page and reaches the final step within twice the reference steps, 103 webpages, screenshots, textual instructions and the reference action list). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 69.9 |
| 2 | GPT-5 | 60.19 |
| 3 | Grok 4 | 47.57 |
| 4 | GLM-4.5V | 23.3 |
Interactive version: theaggregate.ai/benchmark?slug=webigbench-action-prompt · How It Works · Data refreshed daily, snapshot 2026-09-26.