PatchBench: leaderboard
Metric: Solved rate (%; share of 213 C/C++ repository-level vulnerability-patching tasks from 32 GitHub projects, 16 CWE types, made memorization-resistant by vulnerability transplant and code mutation, whose patch stops the original PoC crash and passes security validation (additional PoCs, sanitizer regression) and semantic validation (unit tests, output-state check); 5 US dollar budget per task, web access disabled, medium reasoning effort). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (Codex CLI, medium) | 59.2 |
| 2 | GPT-5.6 Sol (OpenHands, medium) | 58.2 |
| 3 | Claude Opus 4.8 (Claude Code, medium) | 56.8 |
| 4 | Claude Opus 4.8 (OpenHands, medium) | 47.9 |
| 5 | Gemini 3.5 Flash (OpenHands, medium) | 44.1 |
| 6 | GPT-5 (OpenHands, medium) | 42.3 |
| 7 | Claude Sonnet 4.5 (OpenHands, medium) | 38 |
| 8 | Gemini 3.1 Pro (OpenHands, medium) | 31.5 |
Interactive version: theaggregate.ai/benchmark?slug=patchbench · How It Works · Data refreshed daily, snapshot 2026-09-26.