Vul4Py (Direct Prompting): leaderboard

Metric: Plausible patch rate (%; share of the 100 real Python vulnerabilities whose final patch passes the paired oracle: the exploit test must pass and the project's own functional test suite must keep passing; direct prompting: the model gets the vulnerable file, the CVE id and the lines the human fix touched and returns one unified diff, temperature 0). Source: arxiv.org. Saturation forecast: Around 2030. 2 models tracked.

Top models

#ModelScore
1Claude Sonnet 4 (20250514)5
2GPT-4o (2024-08-06)0

Interactive version: theaggregate.ai/benchmark?slug=vul4py-direct-prompting · How It Works · Data refreshed daily, snapshot 2026-09-26.