LongWebBench - Functional Task Success (Single-Image): leaderboard

Metric: Task success rate (%) on LongWebBench's 507 goal-oriented interaction tasks over 129 long webpages generated from one long screenshot of each page: a Gemini-3-Pro actor agent executes each task's steps in a browser on the generated page and a Gemini-3-Pro critic checks the DOM and rendered state; a task succeeds only if every step is executable and the expected outcome is reached; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (Thinking)59.4
2Gemini 3 Pro (Preview)58.38
3GLM-4.6V28.39
4Qwen 3 VL 235B A22B Instruct23.65

Interactive version: theaggregate.ai/benchmark?slug=longwebbench-functional-task-success-single-image · How It Works · Data refreshed daily, snapshot 2026-09-29.