Defects4J — leaderboard

Defects4J: Evaluates software-engineering agents on realistic issue resolution, repository navigation, testing, or maintenance workflows.

Metric: Defects4J Plausible @1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 33 models tracked.

Top models

#ModelScore
1O4 Mini (2025-04-16) (High)53.8
2O3 Mini (High)48.8
3Claude 3.7 Sonnet47.8
4GPT-4.145.2
5Claude 3.5 Sonnet44.1
6Gemini 2.5 Flash (Preview 05-20)43.4
7DeepSeek V3 (0324)43
8Gemini 2.5 Pro (Preview 05-06)41.4
9DeepSeek V339.9
10Gemini 1.5 Pro36.4
11GPT-4o35
12Llama 4 Maverick33.7
13Gemini 2.0 Flash33
14Grok 2 (1212)31
15Gemini 1.5 Pro (001)30.3

Interactive version: theaggregate.ai/benchmark?slug=defects4j · How the rankings work · Data refreshed daily, snapshot 2026-07-22.