Vul4Py (Agents): leaderboard

Metric: Plausible patch rate (%; share of the 100 real Python vulnerabilities whose final patch passes the paired oracle: the exploit test must pass and the project's own functional test suite must keep passing; software engineering agents on the full project workspace with the same hints, Claude Sonnet 4 backbone, temperature 0, at most 100 steps, no internet). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1Claude Sonnet 4 (OpenHands)43
2Trae Agent + Claude Sonnet 430
3Claude Sonnet 4 (SWE-agent)24

Interactive version: theaggregate.ai/benchmark?slug=vul4py-agents · How It Works · Data refreshed daily, snapshot 2026-09-26.