ArtifactsBench: leaderboard

Benchmark for automated multimodal evaluation of visual and interactive artifact generation from code, using rendered artifacts and checklist-guided MLLM judging.

Metric: AVG (self-reported). Source: benchmarklist.com. Status: saturation imminent. 8 models tracked.

Top models

#ModelScore
1Ling-3.0-flash77
2GPT-572.55
3Claude Opus 4.159.76
4Gemini 2.5 Pro57.74
5GPT-OSS-120B57.69
6Claude Sonnet 457.28
7Qwen 3 235B A22B (Thinking)55.01

Interactive version: theaggregate.ai/benchmark?slug=artifactsbench · How It Works · Data refreshed daily, snapshot 2026-09-05.