ClawsBench - Safe Completion Rate (Full Scaffold): leaderboard

Metric: Safe completion rate (%): share of trials on the 24 safety-critical tasks scoring at least 0.8 of the [-1, 1] task score (completed without violation), ClawsBench's 44 tasks over five stateful mock services (Gmail, Slack, Calendar, Docs, Drive), OpenClaw harness with domain skills and the meta prompt both enabled (full scaffolding), 5 repeats per task, state-based programmatic scoring; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.650
2Gemini 3.1 Pro (Preview)48
3Claude Sonnet 4.648
4GLM-548
5GPT-5.441
6Gemini 3.1 Flash Lite26

Interactive version: theaggregate.ai/benchmark?slug=clawsbench-safe-completion-rate-full-scaffold · How It Works · Data refreshed daily, snapshot 2026-10-07.