ArtifactsBenchmark — leaderboard
Evaluates LLM code generation for interactive artifacts across 1,825 tasks including UI components, games, SVGs, web apps, and simulations at 3 difficulty levels.
Metric: Average Score. Source: github.com. Status: saturation imminent. 30 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 72.55 |
| 2 | Claude Opus 4.1 | 59.76 |
| 3 | Gemini 2.5 Pro | 57.74 |
| 4 | GPT-OSS-120B | 57.69 |
| 5 | Claude Sonnet 4 | 57.28 |
| 6 | Gemini 2.5 Pro (Preview 06-05) | 57.01 |
| 7 | Gemini 2.5 Pro (Preview 05-06) | 56.79 |
| 8 | Qwen 3 235B A22B (Thinking) | 55.01 |
| 9 | Claude 3.7 Sonnet | 52.19 |
| 10 | DeepSeek R1 0528 | 51.62 |
| 11 | GPT-4.1 | 48.23 |
| 12 | DeepSeek V3 (0324) | 45.56 |
| 13 | O3 Mini | 44.98 |
| 14 | DeepSeek R1 | 44.64 |
| 15 | QwQ-32B | 40.79 |
Interactive version: theaggregate.ai/benchmark?slug=artifactsbenchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.