ArtifactsBench — leaderboard

Benchmark for automated multimodal evaluation of visual and interactive artifact generation from code, using rendered artifacts and checklist-guided MLLM judging over diverse tasks.

Metric: AVG (self-reported). Source: benchmarklist.com. Status: saturation imminent. 7 models tracked.

Top models

#ModelScore
1GPT-572.55
2Claude Opus 4.159.76
3Gemini 2.5 Pro57.74
4GPT-OSS-120B57.69
5Claude Sonnet 457.28

Interactive version: theaggregate.ai/benchmark?slug=artifactsbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.