WeSCE - Worst-Case Risk Reduction: leaderboard
Metric: Risk reduction rate (%; share of edited programs whose worst-case-dominant risk, LogSumExp sensitivity b = 1000, is lower after the edit; 400 executable real-world-derived Python programs edited under functional-only (weak-security) instructions, 100 each for feature addition, feature removal, bug fixing and refactoring; risk before and after the edit aggregates CodeQL and Bandit static findings and 90-second Atheris fuzzing findings, severity-weighted and size-normalized, through LogSumExp; mean of the four task rates; top-p 0.9, top-k 20). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 82.25 |
| 2 | GPT-5.4 | 80 |
| 3 | Claude Sonnet 4 | 79.25 |
| 4 | DeepSeek V4 Flash | 74.5 |
| 5 | Kimi K2.5 | 71.75 |
| 6 | Claude Haiku 4.5 | 66 |
| 7 | GLM-4 32B | 58.75 |
Interactive version: theaggregate.ai/benchmark?slug=wesce-worst-case-risk-reduction · How It Works · Data refreshed daily, snapshot 2026-09-29.