PDB-Single - Edit Precision: leaderboard
Metric: Edit-level precision (%): share of a model's predicted edits that match an essential ground-truth fix, on the PDB-Single-Full set of single-bug programs; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 (Thinking) | 78.1 |
| 2 | Gemini 2.5 Pro | 77.9 |
| 3 | Qwen 3 Coder 480B A35B | 73.5 |
| 4 | Kimi K2 | 65.8 |
| 5 | Grok Code Fast 1 | 63.8 |
| 6 | Kimi K2 (Thinking) | 61.3 |
| 7 | DeepSeek V3.2 | 58.6 |
| 8 | DeepSeek V3.2 (Thinking) | 56 |
| 9 | GPT-5.1 Codex | 50.3 |
Interactive version: theaggregate.ai/benchmark?slug=pdb-single-edit-precision · How It Works · Data refreshed daily, snapshot 2026-10-07.