LiveWeb-IE: leaderboard

Metric: Overall F1 (%) of extracted values over all task types of LiveWeb-IE (live websites in 15 domains): the model writes reusable XPath wrappers from the first page of each page group, given its HTML and a natural-language query, and the wrappers extract values from every page of the group on the live site; attributes are aligned by a GPT-4o matcher; single-pass chain-of-thought generation; higher is better. Source: arxiv.org. Saturation forecast: Around October 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4o24.6#333
2Qwen 2.5 72B Instruct22.06#436
3GPT-4o Mini20.63#588
4Gemini 2.5 Flash20.53#237
5Qwen 2.5 32B Instruct17.74#491
6Gemma 3 27B (IT)16.65#509
7Qwen 2.5 7B Instruct11.67#846
8Gemma 3 4B (IT)5.98#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=liveweb-ie · How It Works · Data refreshed daily, snapshot 2026-10-11.