LongWebBench - Functional Task Success: leaderboard
Metric: Task success rate (%) on LongWebBench's 507 goal-oriented interaction tasks over 129 long webpages generated from multiple screenshot segments of each page: a Gemini-3-Pro actor agent executes each task's steps in a browser on the generated page and a Gemini-3-Pro critic checks the DOM and rendered state; a task succeeds only if every step is executable and the expected outcome is reached; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 59.24 |
| 2 | Claude Sonnet 4.5 (Thinking) | 55.47 |
| 3 | GLM-4.6V | 30.52 |
| 4 | Qwen 3 VL 235B A22B Instruct | 23.27 |
Interactive version: theaggregate.ai/benchmark?slug=longwebbench-functional-task-success · How It Works · Data refreshed daily, snapshot 2026-09-29.