LongWebBench - Functional Task Success: leaderboard

Metric: Task success rate (%) on LongWebBench's 507 goal-oriented interaction tasks over 129 long webpages generated from multiple screenshot segments of each page: a Gemini-3-Pro actor agent executes each task's steps in a browser on the generated page and a Gemini-3-Pro critic checks the DOM and rendered state; a task succeeds only if every step is executable and the expected outcome is reached; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)59.24
2Claude Sonnet 4.5 (Thinking)55.47
3GLM-4.6V30.52
4Qwen 3 VL 235B A22B Instruct23.27

Interactive version: theaggregate.ai/benchmark?slug=longwebbench-functional-task-success · How It Works · Data refreshed daily, snapshot 2026-09-29.