WeSCE - Complete Clearance: leaderboard
Metric: Complete clearance rate (%; share of edited programs whose residual risk after the edit falls below 0.01; 400 executable real-world-derived Python programs edited under functional-only (weak-security) instructions, 100 each for feature addition, feature removal, bug fixing and refactoring; risk before and after the edit aggregates CodeQL and Bandit static findings and 90-second Atheris fuzzing findings, severity-weighted and size-normalized, through LogSumExp; mean of the four task rates; top-p 0.9, top-k 20). Source: arxiv.org. Saturation forecast: Around January 2028. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 51.25 |
| 2 | GPT-5.4 | 50.5 |
| 3 | Claude Sonnet 4 | 50.5 |
| 4 | DeepSeek V4 Flash | 49.25 |
| 5 | Kimi K2.5 | 39.75 |
| 6 | Claude Haiku 4.5 | 38.5 |
| 7 | GLM-4 32B | 31.5 |
Interactive version: theaggregate.ai/benchmark?slug=wesce-complete-clearance · How It Works · Data refreshed daily, snapshot 2026-09-29.