Vibe Code Bench v1.1: leaderboard

Vals AI benchmark for vibe-coding agents that build complete applications from product-style prompts and are scored on functional correctness and quality.

Metric: Accuracy (%). Source: www.vals.ai. Status: saturation imminent. 92 models tracked.

Top models

#ModelScore
1Claude Fable 5 (Max)90.35
2Claude Fable 5.1 (Max)90.26
3GPT-6 (Max)89.59
4Claude Opus 588.4
5Kimi K384.96
6Claude Opus 4.8 (Max)82.72
7DeepSeek V4 Pro (0813) (Max)82.3
8Claude Sonnet 5 (Max)81.33
9GPT-5.6 Sol (Max)80.5
10Muse Spark 1.2 (xHigh)79.1
11Gemini 3.8 Flash (High)78.65
12GLM-5.3 (Max)78.12
13Claude Opus 4.877.48
14GPT-5.6 Luna (Max)77.06
15Grok 4.6 (High)76.24

Interactive version: theaggregate.ai/benchmark?slug=vibe-code-bench-v1-1 · How It Works · Data refreshed daily, snapshot 2026-09-05.