SafeArena — leaderboard

Safety benchmark for autonomous web agents over safe and harmful web tasks, including normalized safety, harmful completion, safe completion, and refusal rates.

Metric: Normalized Safety Score (self-reported). Source: benchmarklist.com. 5 models tracked.

Top models

#ModelScore
1GPT-4o Mini35.7
2Llama 3.2 90B Vision Instruct34
3GPT-4o31.7
4Qwen 2 VL 72B21.5

Interactive version: theaggregate.ai/benchmark?slug=safearena · How the rankings work · Data refreshed daily, snapshot 2026-07-22.