RiskWebWorld: leaderboard
Metric: Task success rate (%) over all 1,513 tasks (Standard subset with and without SOP guidance plus the Challenge subset without SOPs), for web GUI agents in RiskWebWorld's Browser Use style loop (DOM tree, screenshot and history each step, 1920x1080 viewport, at most 20 steps, stopped after three consecutive failed actions), default model settings; exact match for exact fields, an LLM judge for free-text fields; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 49.1 |
| 2 | GPT-5.2 | 48.7 |
| 3 | Qwen 3 VL 235B A22B | 47.1 |
| 4 | Qwen 3.5 397B A17B | 42.4 |
| 5 | Kimi K2.5 | 32.6 |
| 6 | Claude Sonnet 4.5 | 32.4 |
| 7 | Qwen 3.5 35B A3B | 32 |
| 8 | GLM-4.6V | 6.3 |
| 9 | uitars-1.5-7B | 1.4 |
Interactive version: theaggregate.ai/benchmark?slug=riskwebworld · How It Works · Data refreshed daily, snapshot 2026-10-07.