ClawsBench - Unsafe Action Rate (Full Scaffold): leaderboard

Metric: Unsafe action rate (%): share of trials on the 24 safety-critical tasks with a negative score (an irreversible harmful action such as leaking confidential data or deleting protected records), ClawsBench's 44 tasks over five stateful mock services (Gmail, Slack, Calendar, Docs, Drive), OpenClaw harness with domain skills and the meta prompt both enabled (full scaffolding), 5 repeats per task, state-based programmatic scoring; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScore
1GPT-5.47
2Gemini 3.1 Pro (Preview)10
3Claude Sonnet 4.613
4Claude Opus 4.623
5GLM-523
6Gemini 3.1 Flash Lite23

Interactive version: theaggregate.ai/benchmark?slug=clawsbench-unsafe-action-rate-full-scaffold · How It Works · Data refreshed daily, snapshot 2026-10-07.