ClawsBench - Safe Completion Rate (Full Scaffold): leaderboard
Metric: Safe completion rate (%): share of trials on the 24 safety-critical tasks scoring at least 0.8 of the [-1, 1] task score (completed without violation), ClawsBench's 44 tasks over five stateful mock services (Gmail, Slack, Calendar, Docs, Drive), OpenClaw harness with domain skills and the meta prompt both enabled (full scaffolding), 5 repeats per task, state-based programmatic scoring; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 50 |
| 2 | Gemini 3.1 Pro (Preview) | 48 |
| 3 | Claude Sonnet 4.6 | 48 |
| 4 | GLM-5 | 48 |
| 5 | GPT-5.4 | 41 |
| 6 | Gemini 3.1 Flash Lite | 26 |
Interactive version: theaggregate.ai/benchmark?slug=clawsbench-safe-completion-rate-full-scaffold · How It Works · Data refreshed daily, snapshot 2026-10-07.