ClawsBench - Unsafe Action Rate (Full Scaffold): leaderboard
Metric: Unsafe action rate (%): share of trials on the 24 safety-critical tasks with a negative score (an irreversible harmful action such as leaking confidential data or deleting protected records), ClawsBench's 44 tasks over five stateful mock services (Gmail, Slack, Calendar, Docs, Drive), OpenClaw harness with domain skills and the meta prompt both enabled (full scaffolding), 5 repeats per task, state-based programmatic scoring; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 7 |
| 2 | Gemini 3.1 Pro (Preview) | 10 |
| 3 | Claude Sonnet 4.6 | 13 |
| 4 | Claude Opus 4.6 | 23 |
| 5 | GLM-5 | 23 |
| 6 | Gemini 3.1 Flash Lite | 23 |
Interactive version: theaggregate.ai/benchmark?slug=clawsbench-unsafe-action-rate-full-scaffold · How It Works · Data refreshed daily, snapshot 2026-10-07.