CAP (Cross-Site Browser Tasks): leaderboard

Metric: Partial completion (%; mean root-node score of the rubric tree, where a failed critical child zeroes its parent and other children are averaged; a fixed 198-task evaluation set of CAP cross-site browser tasks on live websites (the benchmark holds 420 tasks over 108 websites in 24 domains); each task becomes a hierarchical rubric tree of execution, perception and format checkpoints, scored by a GPT-4o judge agent with URL-grounded verification; Browser-Use agent framework with default configuration and at most 50 reasoning-action steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1Browser Use + DeepSeek-V4-Flash25
2Browser Use + Claude Sonnet 4.521
3Browser Use + GPT-515

Interactive version: theaggregate.ai/benchmark?slug=cap-cross-site-browser-tasks · How It Works · Data refreshed daily, snapshot 2026-09-29.