CAP (Cross-Site Browser Tasks) - Complex Actions: leaderboard

Metric: Action score (%; mean leaf score over the execution checkpoints for complex UI operations such as map dragging, sliders and video control; a fixed 198-task evaluation set of CAP cross-site browser tasks on live websites (the benchmark holds 420 tasks over 108 websites in 24 domains); each task becomes a hierarchical rubric tree of execution, perception and format checkpoints, scored by a GPT-4o judge agent with URL-grounded verification; Browser-Use agent framework with default configuration and at most 50 reasoning-action steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1Browser Use + DeepSeek-V4-Flash33.7
2Browser Use + Claude Sonnet 4.532
3Browser Use + GPT-529

Interactive version: theaggregate.ai/benchmark?slug=cap-cross-site-browser-tasks-complex-actions · How It Works · Data refreshed daily, snapshot 2026-09-29.