BrowserGym WorkArena L1 — leaderboard
BrowserGym evaluation on WorkArena Level 1: enterprise web agent tasks in a realistic ServiceNow environment. Tests basic workflows.
Metric: Success Rate (%). Source: huggingface.co. Status: saturation imminent. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 79.1 |
| 2 | Claude Sonnet 4 | 63.3 |
| 3 | GPT-5 Mini | 60.6 |
| 4 | Claude 3.5 Sonnet | 56.4 |
| 5 | GPT-OSS-120B | 50.9 |
| 6 | O3 Mini | 48.2 |
| 7 | GPT-4o | 45.5 |
| 8 | Llama 3.1 405B | 43.3 |
| 9 | GPT-5 Nano | 40.6 |
| 10 | GPT-OSS-20B | 38.5 |
| 11 | Llama 3.1 70B | 27.9 |
| 12 | GPT-4o Mini | 27 |
Interactive version: theaggregate.ai/benchmark?slug=browsergym-workarena-l1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.