ProgramBench: leaderboard

Meta and Stanford benchmark testing whether language-model agents can rebuild complete programs from only a compiled binary and documentation. Agents use mini-SWE-agent across 200 open-source program recreation tasks and are scored by hidden behavioral tests.

Metric: Resolved (%). Source: programbench.com. Status: years away from saturation. 21 models tracked.

Top models

#ModelScore
1Claude Opus 5 (xHigh)4.5
2GPT-5.6 Sol (xHigh)1
3GPT-5.6 Sol0.5
4GPT-5.5 (xHigh)0.5
5Gemini 3.6 Flash0.5
6GPT-5.5 (High)0.5
7Claude Sonnet 4.60
8Gemini 3.1 Pro (Preview)0
9GPT-5.50
10GPT-5.40
11GPT-5 Mini0
12Claude Opus 4.70
13Claude Opus 4.60
14Claude Haiku 4.50
15Gemini 3 Flash0

Interactive version: theaggregate.ai/benchmark?slug=programbench · How It Works · Data refreshed daily, snapshot 2026-09-05.