RepairBench — leaderboard

LLM-driven program repair benchmark: 46 frontier models tested on 835 real bugs from Defects4J and GitBug-Java, measuring plausible patch generation at pass@1.

Metric: Plausible@1 (%). Source: repairbench.github.io. Status: saturation imminent. 39 models tracked.

Top models

#ModelScore
1O4 Mini (2025-04-16) (High)50.3
2O3 Mini (2025-01-31) (High)46.4
3DeepSeek R145.2
4Claude 3.7 Sonnet (20250219)44
5Claude 3.5 Sonnet (20241022)41.8
6GPT-4.141.3
7Gemini 2.5 Flash (Preview 05-20)40.6
8DeepSeek V3 (0324)39.6
9Claude 3.5 Sonnet (20240620)39.1
10Gemini 2.5 Pro (Preview 03-25)38.3
11DeepSeek V337.1
12Gemini 1.5 Pro (002)33.2
13GPT-4o (2024-11-20)32.6
14Mistral Medium 332.1
15GPT-4o (2024-08-06)31.7

Interactive version: theaggregate.ai/benchmark?slug=repairbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.