ProgramBench — leaderboard

Meta and Stanford benchmark testing whether language-model agents can rebuild complete programs from only a compiled binary and documentation. Agents use mini-SWE-agent across 200 open-source program recreation tasks and are scored by hidden behavioral tests.

Metric: Resolved (%). Source: programbench.com. Status: years away from saturation. 13 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)0.5
2GPT-5.5 (High)0.5
3Claude Sonnet 4.60
4Gemini 3.1 Pro (Preview)0
5GPT-5 Mini0
6GPT-5.50
7Claude Opus 4.70
8GPT-5.40
9Claude Haiku 4.50
10Claude Opus 4.60
11GPT-5.4 Mini0
12Gemini 3 Flash0
13Claude Opus 4.7 (xHigh)0

Interactive version: theaggregate.ai/benchmark?slug=programbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.