IDE-Bench: leaderboard

IDE coding-agent benchmark with 80 interactive software-engineering tasks that test whether models can make a correct fix on the first attempt.

Metric: Pass@1 Accuracy (self-reported). Source: benchmarklist.com. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.587.5
2GPT-5.285
3Claude Opus 4.583.75
4Claude Haiku 4.578.75
5GPT-5.1 Codex Max73.75
6Qwen 3 Max65
7Gemini 3 Pro55
8Grok 4.1 Fast35
9DeepSeek V331.25
10DeepSeek R120
11Grok Code Fast 111.25
12Llama 4 Maverick2.5
13Llama 4 Scout2.5
14Command-R+0

Interactive version: theaggregate.ai/benchmark?slug=ide-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.