SWE-InfraBench - Passed Tests Share: leaderboard

Metric: Passed tests share (%): share of all unit tests across the tasks that the generated code passes, so partly correct solutions earn credit; 100 AWS CDK (Python) infrastructure-as-code tasks built from 34 open-source and custom repositories: the model writes a masked code block from a natural-language request inside a real project, and the synthesized CloudFormation template is checked by hidden unit tests; one attempt per task, temperature 0.25, Anthropic models prompted with extra XML tags; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 20 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet53.1
2Claude 3.5 Sonnet (20240620)47.9
3Claude 3.5 Sonnet (20241022)47
4Gemini 2.5 Pro (Preview 03-25)42.5
5Gemini 2.5 Pro (Preview 05-06)41.1
6DeepSeek R134.9
7O4 Mini32.1
8O330.8
9GPT-4.126.4
10Llama 3.1 405B Instruct13
11Llama 4 Maverick Instruct13
12Gemini 2.0 Flash Lite11.5
13Claude 3 Haiku10.7
14GPT-4o Mini7
15Qwen 2.5 72B Instruct Turbo2.7

Interactive version: theaggregate.ai/benchmark?slug=swe-infrabench-passed-tests-share · How It Works · Data refreshed daily, snapshot 2026-09-29.