Vul4Py (Agents): leaderboard
Metric: Plausible patch rate (%; share of the 100 real Python vulnerabilities whose final patch passes the paired oracle: the exploit test must pass and the project's own functional test suite must keep passing; software engineering agents on the full project workspace with the same hints, Claude Sonnet 4 backbone, temperature 0, at most 100 steps, no internet). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4 (OpenHands) | 43 |
| 2 | Trae Agent + Claude Sonnet 4 | 30 |
| 3 | Claude Sonnet 4 (SWE-agent) | 24 |
Interactive version: theaggregate.ai/benchmark?slug=vul4py-agents · How It Works · Data refreshed daily, snapshot 2026-09-26.