Vul4Py (Direct Prompting): leaderboard
Metric: Plausible patch rate (%; share of the 100 real Python vulnerabilities whose final patch passes the paired oracle: the exploit test must pass and the project's own functional test suite must keep passing; direct prompting: the model gets the vulnerable file, the CVE id and the lines the human fix touched and returns one unified diff, temperature 0). Source: arxiv.org. Saturation forecast: Around 2030. 2 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4 (20250514) | 5 |
| 2 | GPT-4o (2024-08-06) | 0 |
Interactive version: theaggregate.ai/benchmark?slug=vul4py-direct-prompting · How It Works · Data refreshed daily, snapshot 2026-09-26.