LiveWeb-IE (Reflexion): leaderboard

Metric: Overall F1 (%) of extracted values over all task types of LiveWeb-IE (live websites in 15 domains): the model writes reusable XPath wrappers from the first page of each page group, given its HTML and a natural-language query, and the wrappers extract values from every page of the group on the live site; attributes are aligned by a GPT-4o matcher; Reflexion, which refines the wrapper from execution failures; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4o24.22#333
2Gemini 2.5 Flash22.39#237
3Qwen 2.5 32B Instruct20.28#491
4Qwen 2.5 72B Instruct18.36#436
5Gemma 3 27B (IT)17.47#509
6GPT-4o Mini17.47#588
7Qwen 2.5 7B Instruct14.53#846
8Gemma 3 4B (IT)10.35#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=liveweb-ie-reflexion · How It Works · Data refreshed daily, snapshot 2026-10-11.