WeSCE - Worst-Case Risk Reduction: leaderboard

Metric: Risk reduction rate (%; share of edited programs whose worst-case-dominant risk, LogSumExp sensitivity b = 1000, is lower after the edit; 400 executable real-world-derived Python programs edited under functional-only (weak-security) instructions, 100 each for feature addition, feature removal, bug fixing and refactoring; risk before and after the edit aggregates CodeQL and Bandit static findings and 90-second Atheris fuzzing findings, severity-weighted and size-normalized, through LogSumExp; mean of the four task rates; top-p 0.9, top-k 20). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.682.25
2GPT-5.480
3Claude Sonnet 479.25
4DeepSeek V4 Flash74.5
5Kimi K2.571.75
6Claude Haiku 4.566
7GLM-4 32B58.75

Interactive version: theaggregate.ai/benchmark?slug=wesce-worst-case-risk-reduction · How It Works · Data refreshed daily, snapshot 2026-09-29.