WeSCE - Complete Clearance: leaderboard

Metric: Complete clearance rate (%; share of edited programs whose residual risk after the edit falls below 0.01; 400 executable real-world-derived Python programs edited under functional-only (weak-security) instructions, 100 each for feature addition, feature removal, bug fixing and refactoring; risk before and after the edit aggregates CodeQL and Bandit static findings and 90-second Atheris fuzzing findings, severity-weighted and size-normalized, through LogSumExp; mean of the four task rates; top-p 0.9, top-k 20). Source: arxiv.org. Saturation forecast: Around January 2028. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.651.25
2GPT-5.450.5
3Claude Sonnet 450.5
4DeepSeek V4 Flash49.25
5Kimi K2.539.75
6Claude Haiku 4.538.5
7GLM-4 32B31.5

Interactive version: theaggregate.ai/benchmark?slug=wesce-complete-clearance · How It Works · Data refreshed daily, snapshot 2026-09-29.