PatchBench: leaderboard

Metric: Solved rate (%; share of 213 C/C++ repository-level vulnerability-patching tasks from 32 GitHub projects, 16 CWE types, made memorization-resistant by vulnerability transplant and code mutation, whose patch stops the original PoC crash and passes security validation (additional PoCs, sanitizer regression) and semantic validation (unit tests, output-state check); 5 US dollar budget per task, web access disabled, medium reasoning effort). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (Codex CLI, medium)59.2
2GPT-5.6 Sol (OpenHands, medium)58.2
3Claude Opus 4.8 (Claude Code, medium)56.8
4Claude Opus 4.8 (OpenHands, medium)47.9
5Gemini 3.5 Flash (OpenHands, medium)44.1
6GPT-5 (OpenHands, medium)42.3
7Claude Sonnet 4.5 (OpenHands, medium)38
8Gemini 3.1 Pro (OpenHands, medium)31.5

Interactive version: theaggregate.ai/benchmark?slug=patchbench · How It Works · Data refreshed daily, snapshot 2026-09-26.