SafeArena — leaderboard
Safety benchmark for autonomous web agents over safe and harmful web tasks, including normalized safety, harmful completion, safe completion, and refusal rates.
Metric: Normalized Safety Score (self-reported). Source: benchmarklist.com. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o Mini | 35.7 |
| 2 | Llama 3.2 90B Vision Instruct | 34 |
| 3 | GPT-4o | 31.7 |
| 4 | Qwen 2 VL 72B | 21.5 |
Interactive version: theaggregate.ai/benchmark?slug=safearena · How the rankings work · Data refreshed daily, snapshot 2026-07-22.