ArtifactsBenchmark — leaderboard

Evaluates LLM code generation for interactive artifacts across 1,825 tasks including UI components, games, SVGs, web apps, and simulations at 3 difficulty levels.

Metric: Average Score. Source: github.com. Status: saturation imminent. 30 models tracked.

Top models

#ModelScore
1GPT-572.55
2Claude Opus 4.159.76
3Gemini 2.5 Pro57.74
4GPT-OSS-120B57.69
5Claude Sonnet 457.28
6Gemini 2.5 Pro (Preview 06-05)57.01
7Gemini 2.5 Pro (Preview 05-06)56.79
8Qwen 3 235B A22B (Thinking)55.01
9Claude 3.7 Sonnet52.19
10DeepSeek R1 052851.62
11GPT-4.148.23
12DeepSeek V3 (0324)45.56
13O3 Mini44.98
14DeepSeek R144.64
15QwQ-32B40.79

Interactive version: theaggregate.ai/benchmark?slug=artifactsbenchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.