CAP (Cross-Site Browser Tasks) - Complex Actions: leaderboard
Metric: Action score (%; mean leaf score over the execution checkpoints for complex UI operations such as map dragging, sliders and video control; a fixed 198-task evaluation set of CAP cross-site browser tasks on live websites (the benchmark holds 420 tasks over 108 websites in 24 domains); each task becomes a hierarchical rubric tree of execution, perception and format checkpoints, scored by a GPT-4o judge agent with URL-grounded verification; Browser-Use agent framework with default configuration and at most 50 reasoning-action steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Browser Use + DeepSeek-V4-Flash | 33.7 |
| 2 | Browser Use + Claude Sonnet 4.5 | 32 |
| 3 | Browser Use + GPT-5 | 29 |
Interactive version: theaggregate.ai/benchmark?slug=cap-cross-site-browser-tasks-complex-actions · How It Works · Data refreshed daily, snapshot 2026-09-29.