WebApp1K — leaderboard

React web app code generation benchmark: 1,000 single-task problems across 20 scenarios (blogging, ecommerce, weather, etc.) testing functional correctness with pass@1.

Metric: Pass@1 (%). Source: huggingface.co. Status: saturated. 34 models tracked.

Top models

#ModelScore
1O3 Mini96.1
2O1 Preview95.2
3O1 Mini93.9
4DeepSeek R192.7
5GPT-4o (2024-08-06)88.5
6Claude 3.5 Sonnet (20240620)88.08
7DeepSeek V387.23
8GPT-4o (2024-05-13)87.02
9QwQ-32B87
10Gemini 2.0 Flash (Thinking)85.9
11DeepSeek V2.583.38
12GPT-4o Mini82.71
13Gemini 2.0 Flash82.2
14Mistral Large 2 (Jul)78.04
15Llama 4 Maverick76.9

Interactive version: theaggregate.ai/benchmark?slug=webapp1k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.