SWE-InfraBench: leaderboard

Metric: Correctness (%): share of tasks whose generated code passes all unit tests; 100 AWS CDK (Python) infrastructure-as-code tasks built from 34 open-source and custom repositories: the model writes a masked code block from a natural-language request inside a real project, and the synthesized CloudFormation template is checked by hidden unit tests; one attempt per task, temperature 0.25, Anthropic models prompted with extra XML tags; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 20 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet34
2Claude 3.5 Sonnet (20241022)32
3Claude 3.5 Sonnet (20240620)29
4Gemini 2.5 Pro (Preview 03-25)29
5Gemini 2.5 Pro (Preview 05-06)29
6DeepSeek R124
7O323
8O4 Mini23
9GPT-4.118
10Llama 3.1 405B Instruct9
11Claude 3 Haiku8
12Llama 4 Maverick Instruct8
13Gemini 2.0 Flash Lite5
14GPT-4o Mini4
15Llama 4 Scout Instruct2

Interactive version: theaggregate.ai/benchmark?slug=swe-infrabench · How It Works · Data refreshed daily, snapshot 2026-09-29.