Vellum - LiveCodeBench — leaderboard

Vellum's independently-run LiveCodeBench evaluation. Continuously updated coding problems to avoid data contamination.

Metric: Pass@1 (%). Source: www.vellum.ai. Status: saturation imminent. 23 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro93.5
2DeepSeek V4 Flash91.6
3Kimi K2.585
4Kimi K2 (Thinking)83.1
5Gemini 3 Pro79.7
6Grok 379.4
7Grok 479
8Claude Opus 4.676
9O3 Mini74.1
10Claude Sonnet 4.672.4
11Gemini 2.5 Pro69
12GPT-OSS-120B69
13GPT-OSS-20B69
14DeepSeek R164.3
15Gemini 2.5 Flash63.5

Interactive version: theaggregate.ai/benchmark?slug=vellum-livecodebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.