ProgramBench Almost: leaderboard
Companion ProgramBench metric that counts near-complete program recreations: tasks where the generated implementation passes most hidden behavioral tests but does not fully resolve the benchmark task.
Metric: Almost (%). Source: programbench.com. Status: years away from saturation. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 (xHigh) | 37 |
| 2 | Claude Opus 4.8 (xHigh) | 16.5 |
| 3 | GPT-5.6 Sol (xHigh) | 15.5 |
| 4 | GPT-5.5 (xHigh) | 13.5 |
| 5 | GLM-5.2 | 8.5 |
| 6 | Gemini 3.7 Flash | 5.5 |
| 7 | GPT-5.5 (High) | 5 |
| 8 | Claude Opus 4.7 (xHigh) | 4.5 |
| 9 | Gemini 3.6 Flash | 4 |
| 10 | Claude Opus 4.7 | 3 |
| 11 | Gemini 3.5 Flash | 3 |
| 12 | Claude Opus 4.6 | 2.5 |
| 13 | GPT-5.6 Sol | 2.5 |
| 14 | GPT-5.5 | 1.5 |
| 15 | Claude Sonnet 4.6 | 1 |
Interactive version: theaggregate.ai/benchmark?slug=programbench-almost · How It Works · Data refreshed daily, snapshot 2026-09-05.