WebForge-Bench - Information Retrieval: leaderboard

Metric: Accuracy (%) on the information retrieval and analysis tasks of the 934 WebForge-Bench tasks on generated, self-contained websites with injected real-web noise, up to 50 browser actions in a Chromium GUI, final-state answer checking; multimodal setting: screenshot plus DOM; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScore
1Gemini 3 Pro79.4
2Kimi K2.575.6
3Claude Sonnet 4.573.8
4GPT-5 Mini73.8
5GPT-5.264.4
6Gemini 3 Flash62.5
7Qwen 3 VL 235B A22B58.8
8GPT-5 Nano43.8
9Gemini 2.5 Flash Lite41.9
10qwen3-omni-30B26.2

Interactive version: theaggregate.ai/benchmark?slug=webforge-bench-information-retrieval · How It Works · Data refreshed daily, snapshot 2026-10-07.