ProgramBench Almost — leaderboard

Companion ProgramBench metric that counts near-complete program recreations: tasks where the generated implementation passes most hidden behavioral tests but does not fully resolve the benchmark task.

Metric: Almost (%). Source: programbench.com. Status: years away from saturation. 13 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)13.5
2GPT-5.5 (High)5
3Claude Opus 4.7 (xHigh)4.5
4Claude Opus 4.73
5Claude Opus 4.62.5
6GPT-5.51.5
7Claude Sonnet 4.61
8Gemini 3.1 Pro (Preview)0
9GPT-5 Mini0
10GPT-5.40
11Claude Haiku 4.50
12GPT-5.4 Mini0
13Gemini 3 Flash0

Interactive version: theaggregate.ai/benchmark?slug=programbench-almost · How the rankings work · Data refreshed daily, snapshot 2026-07-22.