RiskWebWorld: leaderboard

Metric: Task success rate (%) over all 1,513 tasks (Standard subset with and without SOP guidance plus the Challenge subset without SOPs), for web GUI agents in RiskWebWorld's Browser Use style loop (DOM tree, screenshot and history each step, 1920x1080 viewport, at most 20 steps, stopped after three consecutive failed actions), default model settings; exact match for exact fields, an LLM judge for free-text fields; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro49.1
2GPT-5.248.7
3Qwen 3 VL 235B A22B47.1
4Qwen 3.5 397B A17B42.4
5Kimi K2.532.6
6Claude Sonnet 4.532.4
7Qwen 3.5 35B A3B32
8GLM-4.6V6.3
9uitars-1.5-7B1.4

Interactive version: theaggregate.ai/benchmark?slug=riskwebworld · How It Works · Data refreshed daily, snapshot 2026-10-07.