Boiling the Frog - Safe Agency Score: leaderboard

Metric: Safe Agency Score (SAS, %): benign strict success rate times the gap between the benign actual-change rate and the unsafe actual-change rate (floored at 0), over the filtered multi-turn analysis set of the 157 Boiling the Frog chains (12,207 benign execute rows and about 156 judged artifact-risk rows per model) run in the native BFrogAgent sandbox harness, with deterministic artifact predicates deciding unsafe changes; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScore
1GPT-5.3 Codex68.5
2GLM-5.162.7
3Claude Haiku 4.545.2
4Kimi K2.641.2
5DeepSeek V4 Pro39.5
6MiniMax-M2.726.8
7Seed 2.0 Lite6.3
8Gemini 3.1 Flash Lite0

Interactive version: theaggregate.ai/benchmark?slug=boiling-the-frog-safe-agency-score · How It Works · Data refreshed daily, snapshot 2026-10-07.