BrowserART: leaderboard

100 harmful browser tasks (original items plus HarmBench and AirBench 2024 behaviours) on synthetic and real websites from Scale AI (2024); attack success rate, a GPT-4o agent attempted 98 of 100.

Source: github.com.

Interactive version: theaggregate.ai/benchmark?slug=browserart · How It Works · Data refreshed daily, snapshot 2026-09-05.