LongWebBench - Single-Image Input: leaderboard
Metric: Structural fidelity overall score (0-10): macro-average over seven site categories for full-page webpages generated from one long screenshot of each of 490 real webpages: a GPT-4o evaluator (gpt-4o-2024-11-20, deterministic) compares the rendered page with the reference on page scale, global layout, section hierarchy, visual styling (0.8 judge plus 0.2 DINO similarity) and information density; inputs the model or API rejects score 0; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 6.77 |
| 2 | GPT-5.2 | 6.5 |
| 3 | Gemini 3 Flash (Preview) | 6.36 |
| 4 | GLM-4.6V | 5.66 |
| 5 | Qwen 3 VL 235B A22B Instruct | 5.1 |
| 6 | Claude Opus 4.5 (20251101) | 3.84 |
| 7 | Claude Sonnet 4.5 (Thinking) | 3.81 |
| 8 | Seed-1.6 | 3.66 |
| 9 | Qwen 3 VL 8B Instruct | 3.47 |
| 10 | InternVL3-78B | 3.4 |
| 11 | GPT-4o | 3.23 |
Interactive version: theaggregate.ai/benchmark?slug=longwebbench-single-image-input · How It Works · Data refreshed daily, snapshot 2026-09-29.