IDE-Bench — leaderboard

IDE coding-agent benchmark with 80 interactive software-engineering tasks that test whether models can make a correct fix on the first attempt.

Metric: Pass@1 Accuracy (self-reported). Source: benchmarklist.com. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.587.5
2GPT-5.285
3Claude Opus 4.583.75
4Claude Haiku 4.578.75
5Grok 4.1 Fast35
6DeepSeek V3.231.25
7DeepSeek R1 052820
8Grok Code Fast 111.25
9Llama 4 Maverick2.5
10Llama 4 Scout2.5

Interactive version: theaggregate.ai/benchmark?slug=ide-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.