ProgramDistill: leaderboard

Microsoft Research's benchmark of reference-guided web development: a coding agent gets a web application with one behavior masked out and a running copy of the original, which it can use but whose source it cannot read, and must restore the behavior in code. Each fix is checked by replaying the recorded browser traces. The board is the public 300-task subset across 26 applications, at restoration depths 1 to 8, all run in one modified R2E-Gym harness. Binary score: the share of tasks whose every replayed behavior passes.

Metric: Binary score (%). Source: microsoft.github.io. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-684.33
2Claude Opus 568.67
3GPT-5.6 Sol60.67
4Grok 4.648.33
5Claude Sonnet 547.33
6GPT-5.3 Codex45.67
7Gemini 3.7 Flash45.33
8Gemini 3.6 Flash29.67
9Gemini 3.1 Pro (Preview)22.33

Interactive version: theaggregate.ai/benchmark?slug=programdistill · How It Works · Data refreshed daily, snapshot 2026-09-25.