TuRTLe Code Completion (Verilator) — leaderboard

TuRTLe leaderboard variant for RTL code completion evaluated with Verilator across VerilogEval MC and VeriGen.

Metric: Aggregated Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 44 models tracked.

Top models

#ModelScore
1GLM-582.23
2Gemma 4 31B80.51
3DeepSeek R1 052878.08
4DeepSeek V3.1 Terminus75.31
5GPT-OSS-120B74.91
6Qwen 3 235B A22B66.8
7GPT-OSS-20B65.92
8Llama 3.1 405B Instruct55.06
9Qwen 2.5 72B Instruct52.29
10Qwen 3 8B48.82
11Qwen 2.5 32B43.61
12QwQ-32B41.14
13starchat2-15B-v0.138.42
14DeepSeek R1 Distill Qwen 14B24.81
15Magistral Small23.47

Interactive version: theaggregate.ai/benchmark?slug=turtle-code-completion-verilator · How the rankings work · Data refreshed daily, snapshot 2026-07-22.