ProgramBench Almost: leaderboard

Companion ProgramBench metric that counts near-complete program recreations: tasks where the generated implementation passes most hidden behavioral tests but does not fully resolve the benchmark task.

Metric: Almost (%). Source: programbench.com. Status: years away from saturation. 21 models tracked.

Top models

#ModelScore
1Claude Opus 5 (xHigh)37
2Claude Opus 4.8 (xHigh)16.5
3GPT-5.6 Sol (xHigh)15.5
4GPT-5.5 (xHigh)13.5
5GLM-5.28.5
6Gemini 3.7 Flash5.5
7GPT-5.5 (High)5
8Claude Opus 4.7 (xHigh)4.5
9Gemini 3.6 Flash4
10Claude Opus 4.73
11Gemini 3.5 Flash3
12Claude Opus 4.62.5
13GPT-5.6 Sol2.5
14GPT-5.51.5
15Claude Sonnet 4.61

Interactive version: theaggregate.ai/benchmark?slug=programbench-almost · How It Works · Data refreshed daily, snapshot 2026-09-05.