WebApp1K Duo — leaderboard
React web app code generation benchmark (hard): dual-task combinations requiring models to implement two features simultaneously, testing compositional code generation.
Metric: Pass@1 (%). Source: huggingface.co. Status: saturation imminent. 49 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 79.38 |
| 2 | Claude Sonnet 4.5 | 78.4 |
| 3 | Claude Opus 4.1 (20250805) | 76.9 |
| 4 | GPT-5.1 | 76.5 |
| 5 | Claude Opus 4 (20250514) | 76.3 |
| 6 | Claude Opus 4.5 (20251101) | 76.3 |
| 7 | Claude 3.5 Sonnet (20241022) | 75 |
| 8 | GPT-OSS-120B | 75 |
| 9 | Gemini 3 Pro (Preview) | 75 |
| 10 | GPT-5 Codex | 75 |
| 11 | GPT-5.1 Codex | 74.3 |
| 12 | O4 Mini | 72.6 |
| 13 | Grok 4.1 Fast (Reasoning) | 71.6 |
| 14 | Kimi K2 (Thinking) | 70.7 |
| 15 | Gemini 2.5 Pro (Preview 06-05) | 70.4 |
Interactive version: theaggregate.ai/benchmark?slug=webapp1k-duo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.